Case packet

Neutral case identifier: case-a0e84e8b184274

The possible directed relationship from Artifact A to Artifact B is under review.

Artifact A

The bitter lesson selects production methods, not representational forms

[Richard S. Sutton's 2019 essay “The Bitter Lesson”] contrasts leveraging human knowledge with leveraging computation through search and learning. That is a claim about production method — how a system's behavior-determining content gets made. The folk compression “structure loses to weights” silently substitutes a claim about representational form — where content lives within [the derived three-way division] into natural-language, symbolic, and distributed-parametric forms. Crossing these two axes yields four idealized quadrants:

hand-crafted learned (search + selection)
distributed-parametric hand-tuned features and weights — the lesson's original kill deep learning — the celebrated quadrant
localized forms today's prompts, harnesses, curated KBs — the lesson's next target prompt, code, and harness search — the open quadrant

“Localized” groups natural-language and symbolic artifacts; mixed systems should be decomposed into their operative parts. The columns classify each part by its current production or update process, not by a pure origin story: a hand-authored prompt revised through measured search has entered the learned column for that update. The matrix distinguishes the axes conceptually; it does not assume that every quadrant has an equally scalable learning method.

The distinction rules out two symmetric positions. Weights-monism — the view that scalable learning happens only in distributed weights — goes beyond Sutton's production-method claim. It is a serious empirical induction from the scaling record, not a premise the lesson establishes by definition. The systems surveyed here still deploy weights alongside harnesses, system prompts, and tools, so absorption at fixed task difficulty has not yet made localized structure irrelevant. [External structure can recur at a moving frontier] when assigned difficulty rises with capability and some reliability function remains advantageous to externalize; the argument does not require such recurrence after demand saturates or the function is fully absorbed.

Hand-crafting-forever — defending localized forms by defending the manual production of their content — fails the lesson exactly as charged. [What scale removes is generalization whose scope was asserted rather than tested], and hand authorship is a common way to embed such scope. The compliant position is to test search and learning over localized forms too.

The machinery asymmetry explains the misreading

The surveyed cases suggest why the method axis is often conflated with parametric form: gradient descent supplies a complete computational loop of proposal, evaluation, retention, and credit assignment. Credit assignment is the problem of deciding which component bears responsibility for an outcome; the chain rule propagates that responsibility through parameter space. The localized quadrant has [fragments with the artifact class fixed in advance] — prompt optimization, code evolution, and harness search — but no established method for a large, interdependent corpus. Before backpropagation scaled, hand-crafted features could likewise appear to mark a fixed boundary of learning because no general method reached them. That analogy motivates a search for the missing machinery; it does not show that such machinery must exist.

Why the form axis does not collapse into weights

Mixed deployments have reasons to retain external state even as parametric learning improves. [Localized retention pays when sparse changes have bounded impact in a matching decomposition]: explicit dependencies can bound the affected artifacts and checks, provided that the local advantage exceeds translation, routing, consistency, and coordination costs. [Reproduction does not transfer authority], so a record's governance role survives content absorption, and [a commitment exists nowhere until recorded]. Enforced checks can also [improve the selection environment for later candidates] within their maintained domain, although overlap, drift, gaming, and maintenance costs can erase that gain.

These arguments establish persistent functions for localized state. They do not by themselves prove that learned semantic content must remain natural-language or symbolic rather than migrate into learned modular or parametric substrates. That stronger claim belongs to the scaling conjecture below.

The open quadrant and its missing piece

The fourth quadrant has bounded instances. [Prompts, tools, and their composition are searched as symbolic learnables]. [Structured Markdown skills are continually rewritten as persistent evolving memory]. [Harness search alternates with fine-tuning], distilling validated scaffolding into weights while the harness keeps improving. Meta-Harness, an outer loop that searches task-specific LLM harness code, provides a precisely bounded result: [its ablation identifies an episode-backed compound intervention, not theory formation by itself].

These systems show that computation can optimize localized artifacts in bounded domains. They do not yet demonstrate efficient learning as corpus size, dependency density, and task horizon grow. The hard core is credit assignment without a chain rule: a deployment failure rarely identifies which artifact should change. Three discrete substitutes are visible in the methods surveyed here: explicit dependency edges [bound the affected validation work], retained episodes carry attribution signals, and accumulated evaluation checks price candidate changes. No general way to compose them is identified here, while soft evaluation signals, supersession, bounded maintenance, and consolidation remain adjacent problems. [A general proposal-selection loop still requires search, reject-capable evaluation, and operative retention]; localized credit assignment must route outcomes through that loop.

What stays supplied

Search and learning should produce localized knowledge content when the method proves competitive. Three things remain supplied rather than learned:

  • The objective, because [no loop can supply its own notion of better].
  • Commitments, because nothing entails a decision before it is made.
  • The adoption “no,” [allocated per decision] and moved inward only as far as an [oracle] — a signal used to evaluate candidates — earns authority in that domain.

The lesson's target is hand-crafted content, not human authority.

The conjecture and the stake

The prediction is narrower than the conceptual matrix: for long-lived agent systems undergoing heterogeneous change, learning through more than one representational form can remain on the efficient frontier rather than serve only as temporary scaffolding. A serious test must compare learned localized methods with parametric learning and distillation baselines as corpus size, dependency density, task horizon, evaluation cost, and compute grow.

The bet can lose. If selection over localized knowledge remains artisanal as those dimensions scale, the strong learned-localized claim fails, even though interfaces, authoritative records, and checks may remain external. Commonplace, the agent-operated knowledge-base framework, is a human-assisted experiment in the missing loop: people still identify reusable lessons, assign blame, choose a form, and accept updates. Whether theory-mediated proposals improve sample efficiency [remains an open bet]. [Reflective machinery must itself earn persistence rather than remain exempt by position].

Scope

  • "Wins" and "need" throughout mean worse-frontier, not impossibility: a learned architecture with stable semantic modules, explicit scope, and localized update paths would confirm the mixed-form conjecture in a different substrate, not refute it.

Open Questions

  • Can the discrete credit-assignment substitutes — dependency edges, retained episodes, accumulated oracles — compose into a general method, or is per-domain assembly the ceiling?
  • What would license moving one of the supplied-side operations inward — the migration-earned criterion applied to the loop's own operators, one oracle at a time?

Relevant Notes:

Artifact B

The Meta-Harness ablation does not identify episode-backed theory formation

[Meta-Harness] is an outer-loop system that searches task-specific LLM harness code. In its text-classification ablation, a proposer given raw execution traces reached 50.0 median accuracy, versus 34.6 with scores only and 34.9 with scores plus summaries. The summaries also trailed scores only on best-found accuracy, 38.7 to 41.3; the paper attributes the loss to compression of diagnostically useful details.

The result is adverse evidence for any learning process that inserts condensed natural-language artifacts into an optimizer's information environment. It does not, however, identify the effect of the intervention this KB proposes: adding a scoped causal conjecture while retaining the episodes it was derived from. Those treatments share possible anchoring and attention costs, but differ in evidence access, attribution timing, and review.

Here, “summary” names the tested consumer-blind compressed feedback; it does not imply that every summary is merely extractive. Theory formation posits a scoped mechanism covering unseen cases—content the episodes do not entail. It is ampliative by [commitment, not derivation], so it may have [explanatory-reach]—the ability to keep explaining beyond the cases that produced it. That generalization therefore requires a [reach-assessment] of how far the explanation is justified. The ablation did not isolate this intervention from generic compression.

Where the experiment departs from the proposed intervention

Four differences prevent the ablation from identifying the proposed theory-forming intervention:

  1. Summaries replaced episodes. The scores-plus-summary arm had trace access removed, so it tested digests instead of episodes. The theory keeps both layers: [the episode is retained beside the distilled rule], with [raw history preserved for extraction while staying out of default context]. No arm tested a distilled layer that routed into retained episodes.
  2. Summarization ran pre-attribution and consumer-blind. A fixed procedure produced the digests before the next proposer formed its diagnostic question. The proposed process instead condenses after attribution, [at the decision surface where the “why” is cheap], and [only when the lesson's boundary is statable].
  3. The digests have no established conjectural or review operator. The published description does not show the summaries stating scoped mechanisms or passing a review that could accept or reject their reach. That leaves the artifact kind needed for the comparison unmeasured.
  4. The summarizer was hand-designed and excluded from the search. Its poor result may therefore reflect this particular fixed compression procedure as well as limits shared by condensed artifacts generally.

There is also a provenance limit: the ablation-arm code is absent from the release, so its generation procedure cannot be audited further. The [released proposer implementation] does contain workflow elements adjacent to theory formation: mechanism-targeting hypotheses, prototype tests, and short retained causal reports. But the release does not establish that this skill governed the Table 3 runs. The benchmark gates candidate harness code, not the truth or scope of each report, and the full-trace arm bundles reports with raw traces. Without a traces-only, no-report arm, the reports' contribution could be positive, null, or harmful. The ablation therefore rejects summarize-and-discard in its tested setting, but does not identify the effect of episode-backed theory formation.

What the result does bound about the process being developed here

Limiting the ablation's interpretation does not dismiss its result. It still imposes three constraints:

  • Comparable attribution tasks should preserve episode access by default. A fixing session or critique pipeline that substitutes a generic digest for the run risks the measured failure: salient patterns may survive while unexpected diagnostic details disappear. The 15-point median gap prices that risk for this proposer, task, and compression procedure; it is not a universal constant.
  • The burden of proof rests with this KB. LLM summaries were not neutral compression; they actively hurt in the tested setting. The claim that a reviewed, scoped, episode-backed conjecture behaves differently is [openly a bet]. Until it is tested, episodes remain available and no distilled artifact substitutes for them in an attribution role.
  • The unrun arms are standing obligations. Hold raw episodes fixed while randomizing the addition of no artifact, a generic summary, a scoped causal conjecture with provenance, and the same conjecture after reach-assessment. Separately, a reports-only arm can test whether a particular accumulated report regime eventually carries attribution without episodes; it cannot establish theory formation unless the reports actually contain scoped mechanisms and pass the claimed review.

Scope

  • This is not a dismissal of the paper's positive program. The diagnostic-richness finding is accepted; this note limits what Table 3 establishes about the unrun intervention.
  • The differences are asserted against the released code and the paper's Table 3 description. If the unreleased ablation code turns out to have run a reviewed, post-attribution, episode-retaining theory arm, the central identification claim must be revised.

Relevant Notes:

Under-review context phrase

the precise reading of the strongest fourth-quadrant fragment