Code-scheduled workflows
Type: reference/types/design-proposal.md
analyse-agentic-system is a multi-step workflow whose schedule lives in prose. The coordinating agent reads the skill, decides what runs next, launches workers, checks their outputs and carries values between steps. Claude Code's dynamic workflows move that schedule into a script, but the script runs only in Claude Code and cannot read files or run commands. This proposal gets as close to a code scheduler as a harness without one permits. A Commonplace program runs the workflow, keeps all run state on disk, and stops only where it needs a sub-agent. The parent agent launches the sub-agents the program names and runs the program again.
Current state (as of 2026-09-29)
- The analysis skill is 518 lines. Its schedule is prose: open the run, freeze sources, runtime baseline, scoping, memory specialist, epistemic lens, reconciliation, set writing, publication. The skill runs in one forked coordinator context (
context: fork). - The memory specialist is the one mandatory fresh worker. The coordinator writes
memory-input.md, hashes it, and launches a worker under Analyse memory and context as the memory analyst withmemory-report.mdas its sole output. The epistemic lens may run locally or in a worker. - Error recovery is the coordinator's. The skill's failure rule tells it to keep a correctable failure in
runningstate, fix the candidate or member, and repeat the failed check. The skill forbids a phase ledger, packet, correction log, retry log or validation receipt. A failed run is not resumed; a new run ID replaces it. - At this dated anchor, mechanical steps were moving into commands. Opening and required route fields were separate proposals; their implemented choices are now workflow-owned opening and route-field acceptance.
- Review jobs already use the worker contract this proposal reuses: code generates each job's
prompt.md, and the worker reads only that prompt and writes one output (ADR 067). The parent dispatches through run review batches, which needs only the ability to launch a sub-agent with a prompt. - Claude Code dynamic workflows exist only in Claude Code. The script sandbox has no filesystem or shell, the resume journal is session-local, and a run takes no mid-run user input.
The problem
- The schedule is interpreted, not executed. Which step comes next, which inputs a worker receives and whether an output is acceptable are decided by the coordinator reading prose. The batch 01 trace audit (2026-09-28, in the analysis-offload workshop's evidence) found costs of this: a step that failed under an exit status of zero followed by validation of a stale report, specialists guessing a wrong contract filename, and correction turns for a report that passed validation.
- The coordinator's context carries the whole run. Every step's state and every worker handoff passes through one conversation. When that conversation ends, the run's progress is recoverable only by re-reading files by hand.
- Code scheduling is not available where the workflow runs. Dynamic workflows supply it in one harness, and their sandbox cannot run
commonplace-*commands, so the deterministic steps would still need the parent conversation. Other supported harnesses have no equivalent.
Expected end state
The design assumes that harnesses will adopt code scheduling of their own, in the manner of dynamic workflows. This is the operator's expectation, not an observed fact. Under it, this design is a bridge, and two requirements follow:
- The workflow definition should read like a native workflow script, so that moving to a native scheduler changes how sub-agents are launched and, where that scheduler can run the definition's language and its mechanical steps, nothing else.
- Machinery that exists only because code cannot launch sub-agents should be as small as possible, because it is the part expected to be discarded.
Vocabulary
These terms are local to this proposal until adoption.
- Workflow definition — the workflow-specific program: the steps, the prompts it renders, the inputs and outputs it declares, and the validators it supplies. It is authored and committed ahead of time, like a saved Claude Code workflow, not generated by the model for each run: publication requires the running package to equal the method commit, so a definition written during a run would fail that check.
- Run — one execution of a workflow definition, owning one run directory. For
analyse-agentic-system, a run is the existing analysis run (AAS-…). - Job — one delegation to one sub-agent: a generated prompt, one output path, and a validator for that output. A step that code can execute is not a job.
- Worker — the sub-agent that executes one job.
- Mechanical step — a step of the definition that code executes and that can be replayed: run again on the same files, it gives the same result and changes nothing further. Examples are assembling a job's input file, building the manifest, and fetching the objects of a commit that is already recorded. It runs again every time the definition is replayed.
- Effect — a step of the definition that code executes and that must run once. There are two reasons a step must run once. It changes the world outside the run, as publishing the finished set does. Or its result cannot be reproduced, as with resolving a branch to a commit, fetching a web page, or reading the clock. Where a step writes is not the test: a web snapshot written into the run directory is an effect, because a second fetch gives different bytes. An effect has a name, declares the files it is made from, and is recorded by code.
- Code orchestrator — the code that executes the workflow definition against a run directory. The agent orchestrator reaches it through two commands, here called
stepandreport. - Agent orchestrator — the parent agent. It runs
step, launches a worker for each jobstepnames, and runsstepagain. It does not decide order or judge outputs. - Report — what the agent orchestrator records, through the
reportcommand, about events only it can observe, such as a launch the harness refused or a repair it made. A report is not a job result.
The model
In the terms of the bounded-context orchestration model, code owns select, transition and the state K. The agent orchestrator implements call_all and nothing else.
step(run directory):
execute the workflow definition from the top
mechanical step -> executed here, at every replay
effect -> executed here, once
agent(prompt, output) -> returns a handle at once
output is accepted: the handle holds the result
otherwise: the job is pending
wait on a handle -> it holds a result: continue
its job is pending: this path stops
no path can continue -> name every pending job and exit
definition finished -> report done and exit
agent orchestrator:
start the run
repeat:
run step
done -> tell the operator and stop
jobs -> read each prompt file and send its unchanged content as one fresh worker's whole message
wait for all of them
when a listed event happens: run report
Four invariants define the model:
- All run state is on disk, and code writes it. The code orchestrator keeps nothing between invocations. It records the jobs it hands out when it returns them. What the agent orchestrator alone observes reaches disk through
report, not through files the agent writes itself. The agent orchestrator keeps only the run's identity, so a fresh session resumes a run by runningstep, once the earlier session's workers have stopped. - No job result passes through the agent orchestrator. Results go from worker to disk to code. The agent orchestrator receives job prompts. It returns only reports, and a report states an observation: code stores it and may show it later, but never accepts an output because of it.
- Control changes hands only at a wait. Calling
agent()does not stop the program. A path stops when it waits on a pending job, andstepreturns only when no path can continue or the definition has finished. Mechanical steps are not shown to the agent orchestrator. - Acceptance is decided by code, and ties one input state to one output. Code records the state of a job's inputs when it hands the job out. The inputs include the prompt, the method files the job depends on, and the launch parameters, so a result produced under one tool scope is not reused under another. An output is accepted only if it passes its validator and the inputs are still in the recorded state, so an output produced from inputs that have since changed is refused. The acceptance names the output's bytes and holds while those bytes and the inputs are unchanged. Otherwise the job is pending again.
A run is started by a separate command. step takes a run and nothing else, so the run must exist before the first step. The agent orchestrator runs one start command with the invocation's arguments, which creates the run directory and returns the run's identity. What the arguments are, and what starting checks, belong to the workflow definition. A reading of outside state that the run depends on, such as the source revision, does not have to happen here: it can be the first effect of the definition, where code records it.
Reports cover listed events. The loop instruction lists the events to report: a launch that failed, a repair, and a stop. A round in which every launch went through needs no report, because code already has its own record of the jobs it handed out and sees their outputs. The list spares the agent orchestrator from judging what counts as abnormal.
Calls are asynchronous. agent() names a job and returns; waiting is a separate act. The jobs launched in one round are therefore every job the program has named and not yet seen accepted when it can go no further. Independent paths need no special construct: two paths that each wait on their own job both stop, and both jobs are returned together. This is the form a dynamic-workflow script has, where agent() returns a promise.
Acceptance is memoization. Asynchronous calls decide when the program stops; they do not carry it across the stop. A stopped program cannot be kept when its process exits, so each invocation replays the definition from the top. A wait on a job whose output is already accepted continues at once, so replay reaches the first unfinished waits. The definition needs no separate readiness function. Replay requires that the definition be deterministic between invocations, that a job's identity not depend on the order in which paths happen to run, and that every mechanical step be safe to meet again. An effect is not replayed; code records it and goes past it. The record alone is not enough: a process can publish and end before it records the publication. When an effect was started and its completion was not recorded, code must establish from the state outside what took place: all of it, none of it, or neither. All of it is recorded and not repeated; none of it is run. Where neither can be established, for example because part of the effect took place, step stops with an outcome stating that the state is uncertain. That outcome is resolved by the operator, who establishes what took place and records it; the agent orchestrator may only stop and report.
The two reasons for an effect differ after an interruption. An effect that changes the world outside may have taken place in part, so establishing what took place means inspecting the state outside. An effect whose result cannot be reproduced writes that result into the run directory, and nothing has used the result before it is recorded. Running it again is then harmless: it counts as completed when the file it writes exists, and as absent otherwise.
A step is split where only part of it must run once. Freezing the sources of an analysis resolves a revision to a commit and fetches that commit's objects. Resolving is an effect, because a later resolution could give another commit. Fetching is a mechanical step, because it reads the recorded commit and gives the same objects each time. Replayed at every step, the fetch restores a checkout that was deleted, and the recorded commit cannot change.
An effect is tied to its inputs, and is not repeated. A job whose inputs change is run again; an effect cannot be treated that way, because running it again publishes twice and skipping it leaves the old result published under a run that reports itself finished. When a completed effect's inputs have changed, code stops and leaves the decision to the operator. The block goes away when the inputs are restored, or when the operator has withdrawn or redone the effect outside and recorded that. The same holds when a later result changes a branch so that the definition no longer calls the effect: replay would skip the call, so code checks every recorded effect when the definition runs to its end, and one that was not reached stops the run until the operator lets it stand or withdraws it.
A step consumes a round. A call of step means every worker launched from the previous step has returned or failed to start. A job handed out before and still without an output counts as a failed attempt. Repeating step gives the same outcome only where nothing is left to consume: on a finished run, on an uncertain effect, and on a run that may only be stopped.
State is on disk because the process cannot wait. The practical scheduler is the host language lets live variables hold K while one process holds the whole run. Here the process exits whenever no path can continue, and the run goes on after it. That is the lifetime mismatch that note names as forcing K into external storage.
Migration. The model differs from a native code scheduler in one place: how a wait on a pending job is satisfied. Here the process exits and a later invocation replays. Where a harness lets code launch sub-agents, the wait is satisfied inside the running process, and the agent orchestrator loop is not needed. The definition is unchanged only if that runtime also runs the definition's language and lets it perform its mechanical steps. The Claude Code sandbox does neither today, so replacing the launch mechanism is the goal of migration and may not be all of it.
Judgment steps are jobs. Scope classification, reconciliation, synthesis and semantic verification currently run in the coordinator's context. Invariant 2 rules that out, so each becomes a job with a generated prompt. The agent orchestrator then holds no analysis state or source content, and prior-analysis exposure in it cannot contaminate the analysis. The cost is that each judgment step needs its inputs declared, which the current skill leaves implicit.
Constraint: minimal core
The shared core contains only what every code-scheduled workflow needs. An element needed by analyse-agentic-system alone goes in its workflow definition. Convergence with other job systems is not a reason to widen the core.
| Element | What it is | Why it is universal |
|---|---|---|
step |
The command that advances the run. It takes the run and nothing else | Code needs one entry point through which it tells the agent orchestrator what to launch. With no other input, its result depends only on the run directory, and the agent orchestrator's call is the same literal text every round. |
report |
The command that records one observation by the agent orchestrator. It does not advance the run | Launching is the one operation code does not perform, so its failures are the one kind of event code cannot see. Without a report they stay in a conversation that a fresh session does not have. |
agent() and wait |
The asynchronous call by which a definition names a job, and the wait by which it asks for the result | Naming a job and needing its result are different moments in any workflow with independent work. Keeping them apart lets step return every pending job at once without the core knowing the dependency structure. |
| Job record | Prompt, output path, declared inputs, validator | The worker receives its whole task in one file, as in ADR 067; code must know where the result lands and how to accept or refuse it. The validator's content is workflow-specific. |
| Input-matched acceptance | Two records: the state of a job's inputs when it was handed out, and the output bytes that passed the validator for that state | Without the first, an output produced from old inputs is accepted against new ones. Without the second, a later write to the output goes unnoticed. With both, resuming means running step. |
| Loop instruction | The agent orchestrator's text | Every workflow is driven this way. With one workflow the text lives in that workflow's skill; a shared snippet becomes worth it when a second workflow repeats it. |
Particular to analyse-agentic-system
These go in its workflow definition, or stay in the commands and skill text it already has:
- the steps and their order (runtime baseline, lens scoping, memory specialist, epistemic lens, reconciliation, set writing, publication);
- prompt templates and input assembly, such as
memory-input.md; - validators, mostly the existing type schemas and
commonplace-validate --fullchecks; - per-job tool scope — what a worker may read, write and run — including the prior-analysis exposure rule and the offload workshop's deny-hook item, stated in each prompt and emitted by
stepas launch parameters for the agent orchestrator to apply where the harness can enforce them; - the start command, which is
openfrom its own proposal; - the finalize endpoint (
commonplace-agentic-analysis-finalize), called by the definition as a mechanical step, and the publish endpoint (commonplace-agentic-analysis-publication), called as an effect; - what the agent orchestrator may touch during a repair, under open choice 1.
The analysis workflow has no operator sign-off step today: invocation authorizes publication. A human-executed step is therefore not in its definition and not in the core.
Open choices
1. How much error recovery the agent orchestrator keeps
Today the coordinator recovers from failures nobody anticipated: it reads the error, investigates and repairs. Code recovers only from failures its author anticipated. In the pure model an unanticipated failure ends the run. The choice is how much of the agent's recovery to keep, and at what cost to its context.
- A. None.
stepretries a failed job within a limit, then ends the run with a failure record for the operator. The loop stays as written above. - B. A blocked outcome. When code cannot proceed,
stepreturns a third outcome. It carries a short statement of what stopped, the path of a failure record on disk, and what the agent orchestrator may do. The agent orchestrator investigates and repairs with its full abilities, then runsstepagain. - C. Repair jobs.
stepturns a failure into a job: a fresh worker receives the failure record and tries the repair. The recovery is still an LLM's, but it costs the agent orchestrator's context nothing and leaves its loop unchanged. A repair job cannot fix what prevents workers from running at all. - D. Both. Repair jobs first, a blocked outcome when they fail or do not apply.
Candidate selection: B, preceded by one code retry that adds the validator's message to the prompt. C is added if blocked rounds are observed to fill the agent orchestrator's context.
Under B, C and D, five rules limit what recovery costs and what it can damage:
- The cost is paid on failure. A round without failure shows the agent orchestrator job prompts and nothing else.
- The outcome points; it does not carry. Evidence stays in the failure record and the run directory. The agent orchestrator reads what it decides it needs.
- The agent repairs; code judges again. After a repair the same validator runs. The agent orchestrator cannot mark an output accepted, so recovery cannot weaken invariant 4.
- The repair scope is declared. The definition states what the agent orchestrator may change. For the analysis workflow the scope is conditions: the environment, a missing input, a misnamed file. The agent orchestrator may also remove a bad output so that its job runs again. It does not write or edit the content of a job's output. Analytical content then comes from workers only. The agent orchestrator is not guarded against prior-analysis exposure, so content it wrote could carry that exposure into the analysis.
- Attempts are counted by code. At the limit, the one permitted action is to stop and report to the operator. The stop ends what the agent orchestrator may do, not the run. The operator can release the stopped job, and the run continues with its accepted outputs.
Two further points hold under every option:
- Workers report problems to disk. A worker that cannot finish writes a problem report to a path beside its output, and replies to the agent orchestrator in one line. Code reads the report on the next
step. - A failed launch is reported, not handled. The output is missing, so
stepnames the job again without being told. The agent orchestrator reports what the harness said when the launch fails, so the reason is on disk ifsteplater returns a blocked outcome, in this session or another. - A repair is reported. After acting on a blocked outcome, the agent orchestrator reports what it changed. The failure record then shows the next reader what was already tried.
- Stopping is reported. An agent orchestrator that stops to hand the run to the operator reports why. It will not run
stepagain, so this observation has no later call to travel on.
2. Where the core lives before a second workflow
- A. Separate core module now.
step,report,agent()and the job record are written as their own module from the start, and the analysis definition uses them. The core can then be tested without the analysis workflow and without a model: small test definitions exercise it, and a test driver plays the agent orchestrator by writing scripted outputs. The module boundary also enforces the minimal-core constraint, because the core cannot import analysis code. The cost is an interface fixed before a real definition has used it. - B. Inside the analysis code, extracted later. The same protocol is written within the analysis package and moved out when a second workflow adopts it. Less code up front; the extraction is later work and the minimal-core boundary is enforced only by review until then. The expected end state favours this option: shared machinery that a native scheduler would replace is cheaper to discard when less of it exists. It also lets the interface be tested before it is fixed: one representative analysis definition, with parallel lenses, reconciliation and a correction cycle, shows whether the core stays as small as the table above claims.
3. Timing against the frozen rerun
The offload workshop decided to fix all backlog items and then rerun batch 01 once, with the method commit frozen before the rerun.
- A. Build before the rerun. The rerun then tests the code orchestrator too, but its results mix one more change into a comparison that already cannot attribute differences to single changes.
- B. Build after the rerun. The rerun measures the offload backlog alone; the code orchestrator is tested on a later batch.
Later options
- Code-side launching where a harness allows it. In Claude Code a dynamic-workflow script could spend one agent call per round on running
stepand returning the job list, then launch the jobs itself. That replaces the agent orchestrator with a script in one harness and tests the migration path, at the cost of an extra agent call per round. - Second workflows. Review jobs and
cp-skill-write-multistageare candidates, not design drivers. Review jobs record completion in the Commonplace store rather than in a run directory, so moving them onto this core would need either a second acceptance back end or a restated review freshness model. Until such a move is proposed, the two share only the ADR 067 worker contract.
Forces
- Recovery versus a clean orchestrator context. Every failure shown to the agent orchestrator adds to its context and invites it to act as coordinator again. Every failure hidden from it must be handled by code that anticipated it, by a repair job, or by the operator.
- Unprompted noticing is lost. A coordinator that reads every output can notice a problem no validator checks. In this model an output is read by an LLM only when a job is assigned to read it. Open choice 1 restores recovery from failures code detects, not detection of failures code misses. The substitute is deliberate: a review job assigned to read the set as a whole.
- Counting attempts needs a record. A retry limit requires failure records on disk, and reports add an event record. The current skill forbids a retry log and forbids resuming a failed run. Adoption replaces both rules; the records are written by code, not kept by the coordinator.
- A report is written by an LLM. It can be missing, wrong or late. Code therefore derives every transition from what it can check — outputs, validators, declared inputs, its own record of jobs handed out — and uses reports for diagnosis and audit. A report that code had to trust would return part of the schedule to the conversation.
- Reporting costs the agent orchestrator attention. A report required every round would be one more thing to get right every round, for information code already has. Reporting listed events keeps a round without failure at one command. An event outside the list is recorded only if the agent orchestrator chooses to report it.
- Two commands versus one. Carrying the report on the
stepcall would keep the core at one command. It would also makesteptake free text, record a report twice when the call is repeated, hold each observation in the conversation until the round ends, and leave no way to record the observation made when stopping. A separatereportcosts one more call in a round that has something to report. - The agent orchestrator is still an LLM loop. Each round costs a parent turn. Because state lives on disk, a fresh agent orchestrator can take over; this is the externalisation recovery named in LLM-mediated schedulers, with the transition logic factored into code and only the launch left in the conversation.
- Launch fidelity cannot be checked by code. Code detects a skipped job, because the output is missing. It cannot detect an agent orchestrator that paraphrased a prompt or did a job in its own context. A fixed launch instruction reduces the risk and does not remove it.
- Acceptance by validator is only as strong as the validator. Most analysis validators check structure. A structurally valid report may still need correction, as the trace audit showed; semantic acceptance is a judgment job.
- Input matching needs declared inputs. A worker that reads an undeclared file makes the match incomplete, so a change to that file does not reopen the job. Declaring inputs per job is also what problem 1 asks for, but it is new authoring work in the definition.
- Replay moves complexity from the definition to the runner. The definition's author writes sequential code. The runner must find every path that cannot continue, keep job identities stable, count retries, and handle effects that may already have taken place. This machinery is the part whose size is not yet known.
- One owner per path. Each job owns its output path and its problem report path. Two jobs that own the same path, or a job that writes into the run's state records, make one job's result look like another's failure; the definition is refused before any job is handed out.
- The operator's acts are not guarded. Releasing a stopped job and recording the state of an effect are meant for the operator. Code cannot tell the operator from the agent orchestrator, so only the loop instruction keeps the agent orchestrator from using them. This is the same limit as launch fidelity.
- One step at a time. Two
stepcalls running on one run at the same moment would each judge and hand out the same jobs. The second is refused while the first runs. - One writer per output. Losing the agent orchestrator's session does not show that its workers stopped. A worker from the earlier session can write to an output path after its replacement starts, and the current skill already forbids a second writer while the first one's ownership is unresolved. The smallest contract is that recovery requires the earlier workers to have stopped. Recovery that tolerates overlap needs a separate output per attempt and a rule for which attempt may be accepted.
- Replay needs discipline. A definition that branches on the clock, or a mechanical step that is not safe to repeat, behaves differently on the second invocation than on the first. Asynchronous paths add one case: a job identified by the order of calls changes identity when paths run in a different order.
- Eager launch. A job is launched once it is named, whether or not the program has waited on it yet. This saves rounds. It also means a job named on a path that a later result makes unnecessary is still run.
- Shared barrier versus staggered progress. The agent orchestrator waits for every worker in a round before running
step. A path whose job finished early does not advance until the slowest job in the round finishes. Advancing it sooner needs the outstanding mark listed under free choices and a more complex loop; the bounded-context orchestration model assumes the shared barrier. - Harness neutrality versus native support. Dynamic workflows give progress display, isolation and concurrency caps in one harness. This design gives them up in exchange for running wherever a command can be run and a sub-agent launched.
- Every new command widens the method paths. Publication requires the running package to equal the method commit, so code orchestrator code joins what must be committed before a run opens.
- Term collision. "Job" already means a review job. The core's job record is compatible with ADR 067 but is not the review store's job; the two stay distinct unless a later proposal merges them.
Free choices
- How acceptance and failure records are laid out in the run directory.
- Whether the definition uses the host language's own asynchronous constructs or explicit handles with a wait call. The first needs a driver that detects when no path can continue.
- What
stepprints per job beyond the prompt path and launch parameters. - The retry limit, and whether a retry's prompt carries the validator's message.
- How
steptreats a job it has handed out whose output is still missing. Under the shared barrier, the nextstepmeans the round is over, so the job is named again. After the agent orchestrator's session is lost, that is correct only once the earlier workers have stopped. - Whether the list of reported events grows beyond the first three, and how much of a report is fixed form rather than free text.
Operativity and warrant
The consumer is the analysis skill. Adoption would replace its step prose with the loop instruction; the judgment content moves into job prompts rendered by the definition. The skill binds the agent orchestrator through the harness's skill loading. Code binds order through what step returns, and binds the workers through the generated prompts. A blocked outcome binds through the same skill text, which states that the agent orchestrator acts only within the scope the outcome names. A report is consumed by code, which stores it, and by whoever later reads a failure record: the agent orchestrator, a repair worker, or the operator. It has no force over acceptance. No consumer exists yet: the skill must be rewritten in the same change that ships the command.
The added automated evaluation is the validator run inside step. For the analysis workflow its warrant is the type schemas and the set checks commonplace-validate --full already runs, covering structure, identity and hash integrity. It does not cover whether a route judgment or synthesis is right; that is assigned to judgment jobs, whose outputs are again checked for structure only.
Adoption criteria
- The operator chooses among the options for choices 1–3, naming the timing against the frozen rerun.
- Each core element still has a universal reason when the analysis definition is written; any element needed only by the analysis workflow has moved into its definition.
- One complete analysis run executes through
stepin each supported harness. On a round without failure the agent orchestrator receives job prompts and launch parameters and nothing else. - An agent orchestrator started in a fresh session, after the earlier session's workers have stopped, resumes a partly finished run by running
step, with no hand-written recovery notes. What the earlier session reported is readable from the run directory. - Before the core's interface is treated as settled, one analysis definition with parallel lenses, reconciliation and a correction cycle has been written against it.
- If choice 2 is A: the core's tests run without the analysis package and without a model.
- A run completes correctly when the agent orchestrator omits every report.
- Tests show that changing a declared input's bytes makes the job pending again, that an output is refused when an input changed between hand-out and completion, that changing an accepted output's bytes makes the job pending again, that an output failing its validator is not accepted, that a second
stepon a finished run reports it finished again, thatstepwithout worker activity uses up an attempt, and that pending jobs on two independent paths are returned in one round. - A test ends the process after publication succeeds and before it is recorded. The next
stepeither recognizes the publication or stops with the uncertain-state outcome; it does not publish again. - If choice 1 is B, C or D: an induced failure that code did not anticipate is repaired without operator help, and the repaired output is accepted by the same validator that a first attempt would face.
Relevant Notes:
- Claude Code dynamic workflows — abstracted-from: the code-coordinates, agents-act division this design copies, with the sandbox, session-local journal and no-mid-run-input limits it avoids
- bounded-context orchestration model — rests-on: the select/call form whose
call_allis the one element left to the agent orchestrator - scheduler-LLM separation exploits an error-correction asymmetry — rests-on: why moving exact schedule state from the coordinator's context into code is expected to remove a class of errors
- the practical scheduler is the host language — rests-on: the definition is host-language code, and run state is reified on disk because the run outlives each process
- 067-Review workers read one prompt and write one output — see-also: the worker contract every job record reuses
- Open analysis runs through workflow start — see-also: the analysis definition's implemented opening operation