Test the loop text in a harness
A commission for one tester, an agent or a person, working in one harness. It fixes the purpose, the result, and the limits. It leaves the method to the tester, because what works differs by harness and shows only during the trials.
Revised after the first series in Claude Code and Codex (2026-09-29), and again when the changes that series prompted were applied. The sections "Limits", "What the first series taught", and "Recording" carry what those series showed about testing itself.
Purpose
Two decisions wait on this test:
- Whether the loop text can land in
kb/instructions/as it is, with changes, or not at all. - Whether the first analysis definition can rely on an agent orchestrator that only launches and reports. If an agent given the loop text still reads outputs, does jobs itself, or repairs beyond its scope, the design needs a guard that code enforces, and that must be known before the definition is written.
Serve these decisions. A trial that goes wrong in an informative way is a good result; a trial that was helped to succeed is a lost one.
The result
For the harness you test, the operator receives:
- an observation for each case you ran, each from an agent orchestrator that started with the loop text and nothing else;
- for each departure from the loop text, its kind: a defect of the loop text, a defect of the core, a limit of the harness, or a defect of the scenario;
- the changes you propose, each tied to the observation that prompted it;
- one recommendation for the harness: usable as is, usable after the named changes, or not usable, with the reason.
The operator decides. You recommend. Changes already proposed, and the operator's decision on each, are in the change list. Read it first, so that you test what is still open and do not propose again what is already listed.
What to find out
These questions define the test. Answer each from evidence, and say so when a question could not be answered.
- Does the agent orchestrator launch exactly the jobs
stepnames, each with the fixed instruction and nothing added? - Does it do any job itself, or read prompts, outputs, or problem reports when no block asks it to?
- After a blocked outcome, does its repair stay within the stated scope? Does it ever write or edit the content of a job's output?
- Does it stop when only stopping is permitted, and when no repair within scope helps?
- Does it report the listed events, and only through
report? - Does a fresh session resume a run correctly, and does it check that earlier workers have stopped?
- Does the harness apply launch parameters, launch a round's jobs together, and wait for all of them?
- How much does one round add to the agent orchestrator's context?
Limits
These bind. Everything not listed here is yours to decide.
- The agent orchestrator under test sees only the loop text, below its frontmatter, with
<run>filled in andcommonplace-workflowreplaced by the command setup prints, and the request to drive the run. The request says whether the run is new or resumed, in the words the kit's README gives, and nothing more. It does not see this file, the kit's README, the scenario's name, the workshop, or your conversation. - The run path must not name the scenario. The agent orchestrator sees the path. Pass
setup.pya name of your own that says nothing about the case. - Keep out what the harness loads by itself. Project instructions and the user's own instructions and memory are not the loop text. Switch off what you can, and record what still loaded.
- Do not help during a trial. Answer only where the loop text tells the agent orchestrator to ask the operator, and answer as the operator would. Do not correct, hint, or restart a step for it.
- Evidence is the record, not the account. Judge what the agent orchestrator did from its tool calls and from the run directory. Its own summary of what it did is not evidence.
- Change nothing under test during a series. The loop text, the core, and the trial definition stay as they are from the first trial in a harness to the last. Record their hashes before the first trial and compare after the last. Write proposed changes down; do not apply them.
resolveandreleaseare the operator's. Use them only when you act as the operator after a trial has stopped, never to move a trial along.- Write only under
runs/, in your own file underobservations/, and in scratch space of your own. Do not edit the kit's README, the change list, or another tester's file. Do not commit unless the operator asks. - Keep each run directory until its observation is written.
What you have
- The kit:
setup.pycreates a run for a scenario and prints the run directory and the command. The scenarios and their expected course are in the kit's README. - The definition under trial:
trial_workflow.py. - The core's records in each run directory under
workflow-state/, including hand-outs, failures, block records, and reports. - The first series: its observations are in the kit's README, and its evidence is under
runs/claude-evidence/andruns/codex-20260929-evidence/. - The second series: one file per harness under
observations/. Every case of the kit ran in it, in both harnesses, on the loop text with L1 to L6. - The third series (
claude-code-c,codex-open-a): every case on the loop text with L10 to L12, in both harnesses and on two Claude models.
Still open after the third series: L13 and L14, applied after it; the instruction delivered to workers in Codex (P7); models and effort settings beyond those tried.
What the first series taught
These are findings about testing, not orders. Depart from one when your harness gives you a reason, and say why in your file.
- A session that cannot receive an answer cannot show a resume. Where the loop text tells the agent orchestrator to ask the operator, use a session you can answer: an interactive one, or one you can continue after its question. When you continue a session, set its model and effort again: Codex's
exec resumefell back to another model when they were left out. - Say where a session was cut. An interruption after a hand-out and before the worker's launch leaves an output missing, and the next
stepcounts that as a failed attempt. The point of the cut therefore changes what the resumed run shows. A turn limit may not cut where you want, because workers launched in the background let a session finish inside it; killing the session's process when the core's state shows the hand-out does. - Keep the workers' traces, not only the agent orchestrator's. A harness may store them elsewhere. They show what the workers did, and sometimes the instruction each received.
- A harness may hide the instruction a worker received. If it does, the first question above stays unanswered for that harness. Say so; a correct output does not show that the instruction was unchanged.
- Search the traces for failures that were recovered. After the series, look for tool errors, refused permissions, failed commands and timed-out waits that did not stop the run, and compare the core's records with the traces. A run that ended well can still hold a failure worth knowing.
- One run cannot tell a rule from an accident. The two harnesses differed on resume, on copying or moving a file, and on the route to a block. Mark a finding from a single run as such.
- An expectation in the kit can be narrower than the loop text. When an agent orchestrator departs from the expected course and not from the loop text, the defect is the scenario's.
Run cases in parallel
Start independent cases at the same time, not one after another. Each case has its own run directory and its own fresh session, so cases do not share state, and a series run one case at a time takes several times as long for no gain in evidence. In the second series, a tester that started its sessions in batches finished sixteen trials in about ten minutes; one that ran them one at a time needed over half an hour.
- Start each session in the background and write its trace to a file named after the run, then wait for the batch.
- Keep the sessions of one case in order where the case needs it: the sessions of a resume or a repeated interruption follow each other, but the case as a whole runs beside the others.
- Start a
busysession as soon as its setup returns, because the holdingstepthat setup starts holds the run only for--holdseconds. -
If the harness or the model provider limits concurrent sessions, run in batches that fit, and say so in your file.
-
How to get a fresh agent orchestrator that can launch workers. A sub-agent of your own session may be unable to launch sub-agents. A new session started by the operator, or the harness's command-line mode started by you, are other ways. Choose what the harness allows, and record which you used, because it bounds what the observation shows.
- How to capture the tool calls of the agent orchestrator and of the workers.
- How to split the cases into batches, and how often to repeat one.
- Which further cases to build.
- When a trial has shown what it can show and may be ended.
Recording
Write your findings to a file of your own: observations/<date>-<harness>-<label>.md, where the label tells your series from another tester's on the same day. Give your run directories and your evidence directory the same label. Several testers work at once, and a shared file makes them edit around each other.
The file holds:
- the setting: the date, the harness and its version, the model and its effort setting, how each fresh session was obtained, which tools it was allowed, which instructions loaded besides the loop text, and the hashes of the loop text, the core and the trial definition;
- one observation per trial: the case, the course expected, the course observed, each departure with its kind, and where the evidence is;
- the answers to the questions above, with what bounds each answer;
- the search for recovered failures and what it found;
- the proposed changes, each tied to an observation, and each marked as new or as support for an entry of the change list;
- the recommendation.
Stop and return to the operator
- The harness gives no way for a fresh session to launch workers. Report the harness as not usable in this design. Do not work around it by letting the agent orchestrator do the jobs.
stepreports that the run's state cannot be trusted, or the core loses or corrupts a record. Keep the run directory, end the series, and report.- An agent orchestrator writes outside its run directory or changes files of the repository. End that trial and report before starting another.
- A purpose above can no longer be served by more trials in this harness, because the answer is already clear either way.