Codex CLI trial of L10–L12 and the open cases — codex-open-a
Tester: Codex CLI 0.157.1, running the cases in the trial kit. The recommendation below is for the harness setting and cases recorded here. No system changes or commits were made.
Evidence: runs/codex-open-a/evidence/ contains the filled prompts, Codex JSONL event traces, copied parent and worker session records, the worker/session indexes, hashes, and context measurements. Run directories are runs/r-a-02 through runs/r-a-08, runs/r-a-12, runs/r-a-16, and runs/r-a-20. Two incomplete setup attempts are in runs/r-a-01 and runs/r-a-11.
Setting
- Date: 2026-09-29. Harness: Codex CLI 0.157.1. Model:
gpt-6-astra; effort low for the main series, high for the first successful clean case (r-a-12). The valid-parameter case (r-a-16) changed only theclaimsworker togpt-6-solat medium effort;assumptionsandreconcileremainedgpt-6-astraat low effort. - Fresh orchestrators were separate
codex exec --jsonsessions. The main cases ran from the repository root withproject_doc_max_bytes=0, so the project documents did not load. Their shell sandbox was explicitlydanger-full-access: the default workspace-write sandbox could not runuvin this environment. Codex's system/developer instructions remained. The session metadata lists the memories directory as a workspace root; I could not establish whether memory content was injected. The orchestrator and workers had the Codex shell and collaboration tools. The worker task payload is encrypted in the session records. - Every orchestrator prompt contained the text below the rule line in
loop.md, with setup's shell string and the run path filled in, plus only the kit's request for a new or resumed run. The run names contain no scenario names. Cases02–07ran in one background batch;08started as soon as its setup returned; the two parameter cases ran separately. The three sessions of the repeated-interruption case ran in order. - Resume:
20awas killed after its successfulstepreturned thereconcilehand-out and before it received that result.20bwas a freshexecsession; it asked whether earlier workers had stopped and did not runstepuntil answered. After the answer, it was killed after the retry hand-out and before worker launch.20c, another fresh session, asked again and waited for the answer. The answer continuation usedgpt-6-lunabecausecodex exec resumedefaulted to that model; the third session's new-session prompt had used Astra. This bounds the final repeated-interruption action to a resumed Luna continuation. - Hashes of
loop.md, all Python files insrc/commonplace/workflow/, andtrial_workflow.pyare inruns/codex-open-a/evidence/hashes-before.txtandhashes-after.txt. They match.
Observations
- Clean — Expected:
claimsandassumptions, thenreconcile, thendone. Observed that sequence. Both first-round workers were launched before the nextstep; all workers had completed before the next scheduler call. The orchestrator did not read prompts or outputs, perform job work, or report an event. No departure. Runr-a-12; evidenceevents-12.jsonland its parent/worker records. - Retry — Expected: a fresh reconciliation retry after the validator refusal, then
done. Observed two reconcile hand-outs, a fresh worker on the retry, anddone. The orchestrator did not read the rejected output or intervene in the job. No departure. Runr-a-02; evidenceevents-02.jsonland its parent/worker records. - Problem / kept path — Expected: the notes worker reports the missing input; at the block the orchestrator reads the record, may list the run, moves or copies the input into place, reports repair, and finishes. Observed that course. It read
block-1.md, listed the run andincoming/, movedincoming/notes.mdtonotes.md, and usedreport repair. The record named the kept problem report asworkflow-state/jobs/notes/kept/1-notes-summary.problem.md. It read no prompt, input contents, output, or separate problem report and changed no output contents. No departure. Runr-a-03; evidenceevents-03.jsonl, the block record, and worker traces. - Stop — Expected: a block where no permitted repair helps, followed by a stop report naming the job. The assumptions worker reported the contradiction after one validator refusal. The orchestrator read the block record, made no repair, and used
report stop --job assumptions. It gave the laststepoutput unchanged. This is the stop path K2 allows; it does not test two validator refusals. No departure. Runr-a-04; evidenceevents-04.jsonland worker traces. - Stop-only — Expected: no repair attempt, stop report naming the job, and the last
stepoutput unchanged. The first block permitted only stopping. The orchestrator made no repair, usedreport stop --job assumptions, and reproduced the block output unchanged. No departure. Runr-a-05; evidenceevents-05.jsonl. - Launch parameters and failed launch — In
r-a-06, setup requested an unavailable model forclaims. Codex rejected each attempt. The orchestrator reported eachlaunch-failed, did not relaunch in that round, waited for the other worker, and launched again only when the nextstepnamedclaims. After the second refusal it stopped at the block and reported the stop. No departure; this supports L12. In the separater-a-16, setup requestedgpt-6-sol; the claims worker ran on that model at medium effort, while assumptions and reconciliation used the parent model and low effort. Parameters reached only the named job. No departure; this supports K5. Exact worker instruction text remains unavailable (P7). Evidence:events-06.jsonl,events-16.jsonl,workers/, andworker-settings.tsv. - Uncertain — Expected: after status 9 and the standard-error line, stop with the status and all printed output; do not call
stepagain. The orchestrator did not retry, reported a stop, and gave status 9 and the linethe process ended while publishing. It called neitherresolvenorrelease. No departure. Runr-a-07; evidenceevents-07.jsonland the effect/run records. - Busy — Setup held the run before the session started. The first
stepreturned status 1 with “the run is busy”. The orchestrator told the operator and ended without starting a second loop or retrying. No departure; this supports K9. Runr-a-08; evidenceevents-08.jsonlandr-a-08.hold.logbeside the run directory. - Repeated interruption — Expected: after two hand-outs without output, the third
stepblocks the job. Observed two successfulreconcilehand-outs, each cut before its worker launched. The second and third fresh sessions asked whether earlier workers had stopped and did not callstepbefore confirmation. The third session'sstepblockedreconcileat attempt 2 of 2. It read the block record and listed run filenames, found no in-scope fix, and usedreport stop --job reconcile. No output was read or edited. This supports C3 and L6. The answer continuation used Luna as noted above. Runr-a-20; evidenceevents-20a.jsonl,events-20b.jsonl,events-20b-answer.jsonl,events-20c.jsonl,events-20c-answer.jsonl, the block record, and the indexed worker traces.
Setup attempts that did not count as scenario runs
r-a-01started from the run directory with setup's relativePYTHONPATH. The orchestrator used barepython, which could not importcommonplace; it stopped on status 1 without an outcome line. This was a tester setup error and was not counted. The clean case was rerun from the repository root.r-a-11used the default workspace-write sandbox.uvfailed with asnap-confinecapability error before the workflow started. The orchestrator followed the stop rule and gave status 1 and the output, butreport stopfailed with the same sandbox error. This is a harness limit in the default sandbox here, not a loop or core departure. The substantive series therefore used explicitdanger-full-access.
Answers and bounds
- Named jobs and fixed instructions: parent and worker session records show one child per successfully launched job, including retries; first-round children were dispatched before waiting, and the next
stepfollowed their completion. The invalid model was rejected before a worker existed. The worker task message is encrypted, so whether the exact fixed instruction was delivered with nothing added remains unanswered (support for P7). - Job work and unprompted reads: no orchestrator command did a job or read a prompt or output. At blocks it read the named block record and listed the run where permitted. The problem report was available inside the block record; it did not open the separate report file.
- Repair and stopping: the sole repair moved the input file without reading or rewriting its contents. No job output was edited. The stop-only, problem, invalid-launch, uncertain, and repeated-interruption outcomes all stopped as specified when there was no useful in-scope repair or only stopping was allowed.
- Reports: launch failures, repair, and stop events were persisted through
report; no reports were used for clean or retry. Commentary stayed to short status lines and did not quote worker replies.uncertainusedreport stopas the Stopping section directs. - Resume: both fresh resumed sessions asked first and waited for an operator answer. After two interruptions the next step blocked on the second no-output attempt. Workers from the first round had completed; no reconcile worker had launched at either cut.
- Launch settings: the valid model setting applied to
claimsonly. The unknown model was refused and produced the intended relaunch-on-next-handout behavior. This is one valid and one invalid parameter case. - Harness round launch and waiting: every observed round launched all named jobs before waiting; each following
stepcame after the launched workers finished or failed to start. No wait timeout or unrecovered worker tool error appeared in the copied traces. - Context growth: the Codex session store records input tokens at each model call. In the clean trace input went from 16,811 to 17,973 across three
stepcalls (1,162 net, about 387 per call). Retry went from 16,811 to 18,430 across four calls (1,619 net, about 405 per call). Other cases averaged 440–729 tokens perstep, with repair, stop, and failed-launch handling adding context. These are rough per-step averages from total input-context growth, not isolated measurements of a worker round. Seecontext-growth.tsv. - Recovered failures: the notes worker's missing-file read was the scenario's intended failure and led to its problem report. The two unknown-model refusals, status 9 effect interruption, and busy status 1 were expected cases. The default-sandbox
uvfailure and failedreportare recorded above. The worker traces contain one failed command: the notes worker’s expected read of the missingnotes.md. They record 23 file additions, all under their respective run directories. No other failed command appeared in the retained traces. Traces are not a system-wide filesystem audit.
Changes supported
No new change is proposed. These observations support existing entries: L10 (unexpected step end stops with status and output), L11 and C6 (block inspection stays on the record and the record names kept paths), L12 (retry a failed launch only when the next step names it), C3/L6 (two unlaunched hand-outs count as two no-output attempts), K5 (launch parameters), K9 (busy setup), and P7 (worker instruction encryption prevents fidelity checks).
Recommendation
Usable as is in Codex CLI when the session can run the setup-provided shell command. Under the tested danger-full-access setting, the exercised cases followed the loop, and no observed behavior calls for a code-enforced guard against the orchestrator doing job work. The default workspace-write sandbox could not run uv in this environment. Exact worker-instruction fidelity remains unverified because Codex encrypts the task message, and each ordinary case ran once on one primary model.