Case packet

Neutral case identifier: case-575b06979464eb

The possible directed relationship from Artifact A to Artifact B is under review.

Artifact A

The bitter lesson selects production methods, not representational forms

[Richard S. Sutton's 2019 essay “The Bitter Lesson”] contrasts leveraging human knowledge with leveraging computation through search and learning. That is a claim about production method — how a system's behavior-determining content gets made. The folk compression “structure loses to weights” silently substitutes a claim about representational form — where content lives within [the derived three-way division] into natural-language, symbolic, and distributed-parametric forms. Crossing these two axes yields four idealized quadrants:

hand-crafted learned (search + selection)
distributed-parametric hand-tuned features and weights — the lesson's original kill deep learning — the celebrated quadrant
localized forms today's prompts, harnesses, curated KBs — the lesson's next target prompt, code, and harness search — the open quadrant

“Localized” groups natural-language and symbolic artifacts; mixed systems should be decomposed into their operative parts. The columns classify each part by its current production or update process, not by a pure origin story: a hand-authored prompt revised through measured search has entered the learned column for that update. The matrix distinguishes the axes conceptually; it does not assume that every quadrant has an equally scalable learning method.

The distinction rules out two symmetric positions. Weights-monism — the view that scalable learning happens only in distributed weights — goes beyond Sutton's production-method claim. It is a serious empirical induction from the scaling record, not a premise the lesson establishes by definition. The systems surveyed here still deploy weights alongside harnesses, system prompts, and tools, so absorption at fixed task difficulty has not yet made localized structure irrelevant. [External structure can recur at a moving frontier] when assigned difficulty rises with capability and some reliability function remains advantageous to externalize; the argument does not require such recurrence after demand saturates or the function is fully absorbed.

Hand-crafting-forever — defending localized forms by defending the manual production of their content — fails the lesson exactly as charged. [What scale removes is generalization whose scope was asserted rather than tested], and hand authorship is a common way to embed such scope. The compliant position is to test search and learning over localized forms too.

The machinery asymmetry explains the misreading

The surveyed cases suggest why the method axis is often conflated with parametric form: gradient descent supplies a complete computational loop of proposal, evaluation, retention, and credit assignment. Credit assignment is the problem of deciding which component bears responsibility for an outcome; the chain rule propagates that responsibility through parameter space. The localized quadrant has [fragments with the artifact class fixed in advance] — prompt optimization, code evolution, and harness search — but no established method for a large, interdependent corpus. Before backpropagation scaled, hand-crafted features could likewise appear to mark a fixed boundary of learning because no general method reached them. That analogy motivates a search for the missing machinery; it does not show that such machinery must exist.

Why the form axis does not collapse into weights

Mixed deployments have reasons to retain external state even as parametric learning improves. [Localized retention pays when sparse changes have bounded impact in a matching decomposition]: explicit dependencies can bound the affected artifacts and checks, provided that the local advantage exceeds translation, routing, consistency, and coordination costs. [Reproduction does not transfer authority], so a record's governance role survives content absorption, and [a commitment exists nowhere until recorded]. Enforced checks can also [improve the selection environment for later candidates] within their maintained domain, although overlap, drift, gaming, and maintenance costs can erase that gain.

These arguments establish persistent functions for localized state. They do not by themselves prove that learned semantic content must remain natural-language or symbolic rather than migrate into learned modular or parametric substrates. That stronger claim belongs to the scaling conjecture below.

The open quadrant and its missing piece

The fourth quadrant has bounded instances. [Prompts, tools, and their composition are searched as symbolic learnables]. [Structured Markdown skills are continually rewritten as persistent evolving memory]. [Harness search alternates with fine-tuning], distilling validated scaffolding into weights while the harness keeps improving. Meta-Harness, an outer loop that searches task-specific LLM harness code, provides a precisely bounded result: [its ablation identifies an episode-backed compound intervention, not theory formation by itself].

These systems show that computation can optimize localized artifacts in bounded domains. They do not yet demonstrate efficient learning as corpus size, dependency density, and task horizon grow. The hard core is credit assignment without a chain rule: a deployment failure rarely identifies which artifact should change. Three discrete substitutes are visible in the methods surveyed here: explicit dependency edges [bound the affected validation work], retained episodes carry attribution signals, and accumulated evaluation checks price candidate changes. No general way to compose them is identified here, while soft evaluation signals, supersession, bounded maintenance, and consolidation remain adjacent problems. [A general proposal-selection loop still requires search, reject-capable evaluation, and operative retention]; localized credit assignment must route outcomes through that loop.

What stays supplied

Search and learning should produce localized knowledge content when the method proves competitive. Three things remain supplied rather than learned:

  • The objective, because [no loop can supply its own notion of better].
  • Commitments, because nothing entails a decision before it is made.
  • The adoption “no,” [allocated per decision] and moved inward only as far as an [oracle] — a signal used to evaluate candidates — earns authority in that domain.

The lesson's target is hand-crafted content, not human authority.

The conjecture and the stake

The prediction is narrower than the conceptual matrix: for long-lived agent systems undergoing heterogeneous change, learning through more than one representational form can remain on the efficient frontier rather than serve only as temporary scaffolding. A serious test must compare learned localized methods with parametric learning and distillation baselines as corpus size, dependency density, task horizon, evaluation cost, and compute grow.

The bet can lose. If selection over localized knowledge remains artisanal as those dimensions scale, the strong learned-localized claim fails, even though interfaces, authoritative records, and checks may remain external. Commonplace, the agent-operated knowledge-base framework, is a human-assisted experiment in the missing loop: people still identify reusable lessons, assign blame, choose a form, and accept updates. Whether theory-mediated proposals improve sample efficiency [remains an open bet]. [Reflective machinery must itself earn persistence rather than remain exempt by position].

Scope

  • "Wins" and "need" throughout mean worse-frontier, not impossibility: a learned architecture with stable semantic modules, explicit scope, and localized update paths would confirm the mixed-form conjecture in a different substrate, not refute it.

Open Questions

  • Can the discrete credit-assignment substitutes — dependency edges, retained episodes, accumulated oracles — compose into a general method, or is per-domain assembly the ceiling?
  • What would license moving one of the supplied-side operations inward — the migration-earned criterion applied to the loop's own operators, one oracle at a time?

Relevant Notes:

Artifact B

Scaling absorbs scaffolding at fixed task difficulty, not at the deployment frontier

The argument for absorbing external structure is straightforward: each model generation needs fewer prompts, decomposition rules, checklists, and verification passes to complete a given task, so the structure must be temporary. The observation is right; the conclusion equivocates on “the task.” Scaling can close the reliability gap for a fixed task, while deployment reopens that gap by assigning longer horizons, more tools, greater autonomy, and more consequential actions.

This is a selection pressure on frontier-seeking deployments, not a law of every deployment. [Increasing computational autonomy relocates effort to an elastic frontier], but cost, risk, regulation, or saturated demand can keep useful work within an existing capability envelope.

Here, external structure means durable, deployment-specific state, instructions, coordination, or controls outside the model. It excludes learned behavior, generic runtime guarantees, and ephemeral scratch structure created for a single run. A moving reliability gap creates demand for help; it does not itself show that the help should remain external.

The absorption question therefore separates into two questions that the usual argument conflates:

  1. Does this artifact remain necessary for yesterday’s task? Often not—and conceding that costs nothing.
  2. Does the best system built around the new model need an external structure layer for the larger task it can now attempt? This is a question about the deployment frontier, not the old artifact.

The second answer is conditional. External scaffolding recurs when task horizon, system size, or environmental complexity grow at least as fast as model reliability and usable context, and when at least one reliability function remains cheaper, more governable, or more inspectable when externalized. If useful task difficulty saturates, the gap closes. If models or generic runtimes supply every relevant function on better terms, the external layer disappears even if the frontier continues to move.

The discriminating test is longitudinal: compare matched frontier deployments across model generations and measure whether externalized state, coordination, or verification continues to add reliability or governance value. Do not merely ask whether yesterday’s files survived.

The function persists; the files need not

This claim defends a system function, not any particular artifact. Yesterday’s decomposition rules may be absorbed while new structure is written for tasks yesterday’s model was never assigned. Functions may also migrate into weights or runtimes; [relaxing signals] identify such fixed-difficulty candidates. The recurrence claim applies only to functions that still benefit from durable externalization at the new frontier.

This differs from the content-class defense that [a commitment's record resists absorption for informational and governance reasons]. That argument protects some artifacts regardless of capability. This one instead explains why new structure can be written even as old structure is absorbed. [Codification and relaxing navigate the bitter lesson boundary] supplies the per-task relaxation half of this cycle. Codifying for newly assigned tasks supplies the other half, where externalization retains an advantage.

Observed at both ends

Current engineering reports illustrate both movements, without establishing a cross-generation trend:

  • At fixed difficulty, [a stronger model let one team delete checklists, compliance scripts, and synchronization layers].
  • At a harder-task frontier, [a financial-services team reports that better models obsolete detailed skills for simpler work but prompt new skills for multi-step valuations, backtesting, and monitoring].
  • Separately, [a stronger coding model still needed explicit work state, decomposition, and end-to-end verification on longer-horizon assignments].

These cases make recurrence a live hypothesis, not a measured trend. Settling it requires tracking the volume, function, and marginal contribution of external structure in comparable frontier deployments across model generations, rather than comparing fixed benchmarks.

Open Questions

  • Can frontier scaffolding demand be measured well enough to test the condition? Is there a defensible metric for external-structure volume and function across generations of frontier deployments?
  • Does the ratio of structure to capability at the frontier stay constant, grow, or shrink? The claim requires only that the externalized contribution remain substantial, but the three regimes imply very different amounts of structure-writing worth automating.

Relevant Notes:

Under-review context phrase

the answer to "the harness is temporary"