MerchantBench: ReAct and optional persistent notes

Type: types/note.md

This analysis covers the reference ReAct baseline, its SDK and consumed server admission, observation, memory and trace interfaces. It excludes other agents, simulator market algorithms, private data, evaluation launchers and provider internals. Findings are static and code-grounded at the pinned commit; no benchmark or model was run.

The runtime receives a simulated store observation, asks its configured model for tool calls, submits them to the environment and feeds results into later model invocations. No-call responses and exhausted hop budgets force end_of_step. HTTP410 or a configured operating horizon ends the client. The environment supplies feedback and an outcome metric, while this loop implements no independent oracle for the best business action. Baseline loop.

Its 160k-to-30k context maintenance is approximate recent-history trimming. When the write tool is exposed, a reminder invites the model to preserve important details first; trimming follows any HTTP-successful action, including a failed memory write returned as a tool result. With the tool denied, history is trimmed before inference. A single large last message can exceed the nominal history budget, and model tools/system overhead is separate. Maintenance code, trim helper.

The pinned default enables read_memory_doc and write_memory_doc: their denylist entries are commented out, contrary to the agent README. The optional path lets a model derive run-local Markdown strategy or follow-up notes and request them after raw history has disappeared. It affords per-task online trace learning, without guaranteeing a write, recall or benefit. The current document is overwritten under a 256KiB cap, then history is appended; these are separate file operations. Only current content has the supported read tool. Default configuration, scratchpad implementation.

Server controls check step freshness when a header is supplied, scenario permissions, argument shape, hook state and quota. The supplied SDK sends the step header. Mutating-call identities/fingerprints support bounded replay control; sequential batch execution and later trace persistence do not establish an atomic whole-turn transaction. Authentication is configuration-dependent. Local stale-step recovery slices history using an old length, which can be weakened by intervening trimming; it does not undo environment effects. Admission path, optional authentication.

The strongest supported contribution is a model-action-feedback loop with optional reusable notes and inspectable protocol controls. Prompt-size and failure-state feedback support narrow reflection on runtime state. Conjectural learning and self-improvement remain uninspected: writable strategy text and simulator feedback do not establish criticism of an operative theory with improved future capacity. The exact result retains quotes, branch-specific memory comparison and epistemic limits. Controlled note-recall interventions and fault/replay traces would strengthen separate conclusions.