Ingest: Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Type: types/ingest-report.md

Classification

A research preprint combining a language-level account of recursive agent calls, a framework description, and controlled benchmark comparisons. Zhening Li and coauthors are affiliated with MIT CSAIL and independent research; the authors developed JAZ and conducted its evaluation. The source supplies experimental methods and individual run results, rather than an independent replication.

Summary

JAZ exposes one agent primitive, invoke, whose body is generated by a model and executed in a Python REPL. Calls can recursively invoke agents, pass ordinary objects and callables, and return ordinary objects. Its distinguishing addition to code-mode agents is programmatic access to the user prompt and interaction history, so agents can pass instructions and traces without regenerating their text. On StuLife, GPT-5.4 nano achieves 69.9% far-recall pass rate, versus 32.0% for the authors' CodeAct-with-subagents ablation and 61.8% for their configured Letta baseline. The ablation removes both prompt and history variables, with corresponding prompt adaptations; it does not isolate either variable alone. On AppWorld, a GPT-5.4 meta-agent revises nano solvers' prompts and tools across tasks, achieving 74.2% task-goal completion versus 69.9% for ACE, whose adaptation is capped at the first 42 of 417 tasks. These results support the sufficiency of the supplied recursive REPL and guidance for two workflows, within the tested models, interfaces, and adaptation schedules; they do not establish that specialized harnesses are generally unnecessary.

Contribution assessment (our opinion)

The useful contribution appears to be a small, testable extension of code-mode agency: make the agent's own instructions and continuing history available as objects, then use the same returning call for continuation and adaptation. That targets a real source of loss and overhead when models must reconstruct material already present in their context. The formalization clarifies composition, while the controlled JAZ ablations give stronger evidence than comparisons between unrelated implementations alone. Likely significance lies in exposing an interface choice that can be tested elsewhere. The evidence is narrower than a general case for maximal expressivity: two benchmark regimes, joint removal of two facilities, and a restricted ACE schedule leave the universal architectural claim open.

Quotes

However, every delegation preserves the full history completely through the prev_history, which has accumulated every prior agent’s history. The agent spent one turn searching all protocol names in its prev_history. All the relevant lectures were correctly retrieved, and the agent answered correctly in its next turn. --- kb/sources/.snapshots/harness-as-a-language-jaz.md @ sha256:e5421e433ba0c6a06a1af7debee98a1419160d128edab016dfe6bb175d0eb4f7 — Section 4.1, Analysis, p. 7: JAZ on StuLife task 1282.

Letta Agent used its conversation_search tool to search through its conversation history for exact protocol names mentioned in the question. Using a combination of keyword search and vector search, the tool returned the top-ranked hits, which were either the quiz question itself, or earlier messages about similarly named but different protocols. In this scenario where the agent needs exact substring matching, the only tool Letta had (conversation_search) did not support it. --- kb/sources/.snapshots/harness-as-a-language-jaz.md @ sha256:e5421e433ba0c6a06a1af7debee98a1419160d128edab016dfe6bb175d0eb4f7 — Section 4.1, Analysis, p. 7: Letta on the same recall question.

We apply the CodeAct hook that removes history and the user prompt(s) from the REPL of every invoke, and make the minimal modification to the long-horizon guidance and ContextWindowWarning prompt to accommodate the removal. --- kb/sources/.snapshots/harness-as-a-language-jaz.md @ sha256:e5421e433ba0c6a06a1af7debee98a1419160d128edab016dfe6bb175d0eb4f7 — Appendix C.2.2, Method Details, p. 16: joint removal and guidance adaptation.

Letta provides two search backends. The SQL backend caused most search results to turn up empty: it applies exact substring search to the entire search query, yet the agent issues search queries that are more appropriate for conventional search engines and thus frequently do not appear as an exact substring of anything in the history. We thus used Letta’s more powerful search backend, Turbopuffer, which implements hybrid search, combining BM25 keyword-based search and vector-embedding-based search. --- kb/sources/.snapshots/harness-as-a-language-jaz.md @ sha256:e5421e433ba0c6a06a1af7debee98a1419160d128edab016dfe6bb175d0eb4f7 — Appendix C.2.2, Method Details, p. 16: SQL query mismatch and choice of hybrid backend.

Connections Found

The paper supplies a bounded empirical case for the context-operation interface claim. JAZ can search retained REPL entries directly and pass selected material into the next call. In the authors' worked StuLife example, the configured Letta hybrid search returns irrelevant hits while JAZ finds the needed lecture by substring. This supports examining the available operations and actual projections separately. It is not a matched comparison of retrieval operations alone: the systems also differ in capture, prompting, and execution, while the closer JAZ ablation jointly removes prompt and history access.

The source also supplies an architectural example for the host language as practical scheduler: returning recursive calls compose through ordinary control flow, and live variables carry state. The experiments exercise that arrangement within running processes. Their success does not remove the note's lifetime and capacity boundaries, even though the framework separately describes replay hooks.

For learning, the source makes the fixed-decomposition boundary concrete. The meta-agent can revise solver inputs and batch sizes, but the experiment fixes the models, environment APIs, feedback, task order, and meta-agent/solver split. The gain over the restricted ACE configuration compares compound adaptation arrangements inside those boundaries.

Learning Claims (our opinion)

JAZ's continual self-improvement changes external prompts and executable tools, not model weights. A meta-agent runs task batches, reads returned solver histories and test reports, diagnoses failures, and modifies inputs for later tasks. Appendix E.2 instructs it to state a hypothesis about the cause, make a minimal targeted change, and discard changes that fail subsequent validation. Our mapping to theory building is therefore about the whole meta-agent/solver process, with different evidence strength for each condition:

  • Localized content: prompts and named Python functions are explicit revision targets. Their existence is clear from the described mechanism; the paper does not provide a detailed retained inventory of the resulting hypotheses or their scope conditions.
  • Consumption: revised prompts and tools are passed to later solver calls, and the authors report that this happened. This establishes a described consumption path, without isolating the behavioral effect of each content change.
  • Content-directed criticism: the supplied guidance explicitly requires causal hypotheses and targeted tests; the qualitative analysis reports investigation of failure traces. This supports the intended process, but the paper does not expose enough actual hypothesis–criticism–revision sequences to verify how consistently criticism addressed identified content rather than merely reacting to scores.
  • Iteration: the authors report repeated updates and changing batch sizes across the task sequence. Criticism-informed changes can persist across tasks within one run; cross-run or cross-problem reuse is not demonstrated.

The design thus fits a theory-builder interpretation, with its criticism condition less directly evidenced than its explicit state and iteration. Prompt and function parts permit selective repair in the sense of addressable theory, but the study does not test the benefit of finer addressability. It adds a case where such a process can operate in live REPL state without a durable library; it does not establish that this persistence horizon is sufficient for Commonplace's purposes.

Improved capacity for future action is a separate claim. The reported average exceeds the independent per-task CodeAct baseline (74.2% versus 67.5%), but that comparison adds a stronger meta-agent as well as adaptation. There is no matched frozen-input meta-agent control that isolates the effect of revising content. All adaptation receives full benchmark test feedback under one shuffled task order. The result supports usefulness of the compound online arrangement on that sequence, with neither transfer to another problem distribution nor the adequacy of its fixed decomposition established.

Extractable Value

  1. Treat instructions and traces as objects that can be passed intact. Within the tested code-mode setup, joint access to prompt and history variables reduces reliance on model-generated reconstruction. This gives the existing context-interface note a concrete failure mechanism and comparison; it does not identify history access alone as the cause. [quick-win]
  2. Separate trace access from retained lessons. Exact history lets a meta-agent selectively inspect failures; prompts and functions carry its attempted corrections into later tasks. The distinction helps describe both the evidence supplied to criticism and what later behavior consumes, without equating trace retention with learning. [quick-win]
  3. Test adaptation schedules alongside update targets. JAZ can batch tasks and modify both prompts and tools, whereas the evaluated ACE adapts its playbook after each of only the first 42 tasks. The measured cost and quality difference motivates a matched comparison of schedule and update targets, not a conclusion that one named learning architecture generally dominates. [experiment]

Limitations (our opinion)

The strongest controlled contrast adds prompt and history access together, plus the necessary guidance changes. It does not distinguish their contributions or compare against an otherwise matched agent given a narrower but reliable trace-access tool. The worked recall example diagnoses one failure, not all of the benchmark gap. Letta's SQL backend supports exact substring queries, but the authors selected its hybrid backend after unsuitable query behavior made SQL retrieval ineffective; the reported search failure should not be generalized to every Letta configuration.

The AppWorld advantage over ACE is 4.3 percentage points with six runs per adaptive method; reported standard errors are 2.1 points for JAZ and 1.4 for ACE. One JAZ run scores below the nonadaptive CodeAct mean. ACE's capped learning period, removed environment-specific guidance, and changed reflector/curator model make this a comparison with the paper's particular adaptation setup. The large gap over official AppWorld CodeAct is especially unsuitable as evidence for self-improvement: the authors attribute most of it to ambiguous submission instructions that their own baseline fixes.

The prompt-only designation still includes substantial supplied method guidance, environment tools, monitoring hooks, and completion guards. StuLife uses adapted tools and an explicit continuation template triggered at 70% context use. AppWorld fixes a two-level call structure and supplies full test feedback. As the fixed-decomposition note cautions, improvement inside these choices does not validate alternatives that were never compared. No implementation was inspected or executed for this ingest. The source does not establish cross-restart memory reliability, production-scale retention costs, or improvements without benchmark feedback.

Update the context-operation interface note with JAZ as a bounded empirical example, retaining the joint prompt/history ablation and the configured-search limitation beside the recall result.