ARC skill as a route-asymmetric epistemic architecture

ARC is a route-asymmetric epistemic architecture. Its checks address different targets at different times and exert different force. The sixth-case comparison must therefore classify authority routes and report epistemic authority separately from operational authority, rather than assign ARC one system-level oracle or one participation/containment cell.

An authority route is a checked target together with the oracle that checks it, the timing of the check, its implemented force, what the result licenses a consumer to believe, and how much later behavior the result permits. Here, epistemic authority means what may be relied on as warranted. Operational authority means how much behavior an artifact may authorize before another check. An oracle is operative only when its result changes admission, survival, rollback, use, or continued execution. Recording a grade is not enough.

This separation matters because ARC has no single object that can safely be called “the model.” Its relevant objects include pre-action prediction text, individually parsed prediction claims, an optional ungraded reason, prose NOTES.md, an executable rules.py model, a generated plan, environment progress, and a campaign scorecard. These objects have different parsers, consumers, checks, and effects. A statement that “the model passed” is ill-typed until it names the object and route. The same point follows from Commonplace's existing work on artifact-analysis axes, behavioral authority, action-model consumption paths, and typed verification targets; ARC is a worked application of that path-level analysis, not evidence that ARC alone discovered a universal taxonomy.

This reading inspects ARC at commit dba53c3. The local related-systems/ checkout is ignored by the published repository, so every ARC link below is pinned to that immutable revision. The evidence order is route-specific. The selected implementation is strongest for what the harness requires, grades, records, refuses, or discards: the prediction parser and grader, live execution, executable-rule replay and search, inspection and status, event core, and CLI routes. The ARC skill doctrine and ARC repository README establish instructions and author claims, but not agent compliance. Campaign quantities come only from the README because this checkout supplied no run directories. It would therefore be wrong to compress the architecture to “ARC verifies theories before acting”: before a manual action, ARC checks prediction admission; consequence grading, executable replay, plan freshness, and task outcome check other targets at other times.

Authority-route ledger

Target / route Oracle and timing Implemented force Epistemic authority Operational authority
Manual-action admission Before the action, the CLI requires a nonempty, parseable --predict before live execution reaches the environment. The accepted grammar is defined by the prediction parser and grader and exposed through the CLI routes. An absent or malformed prediction refuses the command. It establishes pre-registration in the accepted grammar, not prediction quality, causal discrimination, or explanation quality. The parser is deterministic for this narrow admission target. It does not discriminate whether admitted text meets the README's broader “falsifiable claim” target or carries causal or explanatory quality; those are different targets, not one unqualified oracle-strength ordering. The manual action is either admitted or prevented before it reaches the environment.
Structured prediction claim After the paid action settles, the after-frame, level count, or game state grades each parsed clause. The event stores per-clause grades and aggregate prediction success or surprise. A result bears on the stated observable consequence on that one transition. Under fine-grained theory warrant, it does not transfer automatically to a prose reason, unique causal model, executable rules, strategy, or benchmark result. A standalone miss is recorded, but it creates no persistent block on a later standalone action.
Manual batch Every step has a parsed prediction. After each executed action, live execution checks the result. The first miss, level advance, game over, or win discards the remaining queue suffix. Each grade retains the same transition-bounded meaning as a standalone prediction grade. The surprising action has already happened, but no later action in that queue executes. This is a one-surprise forward exposure boundary over one active queue, not rollback or theory rejection. The scope follows the path-and-force distinctions in operative change and behavioral authority.
Retained event evidence and inspection After actions, the event core retains a JSONL timeline with settled frames, actions, environment state, and level state; live execution adds prediction fields and receipts. Inspection and status exposes grids, diffs, prior events, and explicitly labeled click hypotheses. Episodes remain inspectable and available to later note repair or executable replay. Retention supports checking and re-derivation; it does not by itself accept a theory. This is the distinction made by episode/rule re-derivability. The evidence can inform later actions and checks, but mere storage grants no action route on its own.
Prose notes and level transfer The ARC skill doctrine asks agents to separate Verified, Assumed, and Plan, cite event IDs, repair misses, and re-test prior mechanics after a level change. The harness initializes, archives, displays, warns, and nudges through CLI routes, live execution, and inspection and status. No selected action route semantically validates headings, citations, repairs, or re-tests. A later file edit can clear the modification-time level-transfer warning; repeated misses produce a nudge, not a gate. Note status is self-declared. Event citations may preserve lineage when followed, but the harness does not grant warrant to the prose labels. This is doctrine with partial methodology enforcement. Notes are displayed and available to guide later actions if an agent consumes them; the supplied code and absent run evidence do not establish realized influence. No note-level accept/reject transition constrains that possible use.
Executable-model replay Offline, executable-rule replay and search re-grounds at the opener, resets, unmodeled transitions, and level boundaries. Ordinary nonterminal transitions compare exact render output when supplied or equality under the author's observe projection. Level completion instead checks goal(predicted); game over checks the optional dead(predicted) predicate. Replay returns MISMATCH, INCOMPLETE, or HISTORY_FIT. HISTORY_FIT means replay found neither a mismatch nor a declared gap under the check applicable to each transition. It does not establish unseen mechanics, explanation quality, transfer, or global truth. INCOMPLETE records explicit gaps rather than falsehood. These limits follow from fine-grained theory warrant, explanatory-reach, and formal-system explanatory-reach. A MISMATCH blocks rule-model search. Replay status does not otherwise execute environment actions.
Rule-model search After replay, A* searches the executable model's actions and step relations. It skips Unknown edges and uses goals defined by the model. MISMATCH categorically refuses search. INCOMPLETE may proceed only if replay leaves a usable current model state; otherwise search refuses because the current board cannot be grounded. A found path establishes reachability inside the supplied modeled region, not reachability in unmodeled environment states. A locally grounded model may produce a plan even while other recorded history remains explicitly incomplete.
Generated-plan applicability and live execution Before execution, ARC compares the plan's source event, current observation hash, and rules.py hash with current state. After each action, it compares the result with the saved prediction. Stale provenance refuses the plan. First live divergence discards the remaining suffix. Matching provenance establishes current applicability, not endorsement of the producing model. Encountered live fit supports only the checked plan step along the executed route. Plan use is tightly bounded: unchanged inputs permit an attempt, and live divergence stops later queued actions. The first divergent action is not rolled back. This is strong queue-scoped path control, not general model acceptance.
Environment progress After an action, returned environment state and completed-level count determine progress. Level or game completion controls the task transition, including when the same action's prediction misses. It establishes task progress or completion, not prediction fit or model warrant. The environment ends or advances the relevant task state.
Campaign scorecard After whole runs, the ARC repository README reports benchmark replay and aggregate outcomes. No supplied implementation route turns that score into theory status. It is an external, whole-system reported result rather than an in-harness acceptance transition. At most it bears on the configured bundle under the reported setup. It does not identify a component contribution, by the experimental contrast limit and fine-grained theory warrant. The environment outcome has already ended the run; the scorecard itself grants no additional action authority in the selected harness.

Grading supplies evidence on every graded action, but it acquires code-enforced continuation force only when a queued route consumes the result. Operativity otherwise lands on admission, model-search eligibility, plan applicability, queue survival, or environment completion. “The operative oracle” therefore has no unqualified system-wide referent in ARC.

Where epistemic and operational authority diverge

Prediction admission, consequence fit, and queue survival

ARC's pre-action gate deterministically enforces a narrow admission target. Unrecognized prose is stored as an ungraded note; only when the prediction contains no structured gradable claim does the parser append a generic change claim. A cell claim checks the after-frame value at one coordinate and can pass when that value was already present. The optional --because reason is recorded but not graded. These examples do not show that every parsed clause is coarse or that campaign agents deliberately chose weak claims. They show that admission does not require a discriminating causal explanation. (Prediction parser and grader; live execution; CLI routes.)

The ledger separates three transitions: grammar controls admission; observation grades each structured consequence after action; and only a queued route consumes a miss to stop later actions. The first mismatching queued action is spent, only that active queue's suffix is prevented, and a later standalone action remains available. A hit or miss still bears on one parsed consequence on one transition. Forward containment neither identifies nor rejects the producing theory, because fine-grained theory warrant and explanatory-reach do not expand with the result's operational force.

Notes and executable rules are unequal episode/representation pairs

ARC retains episode evidence beside prose notes and executable rules. The note relation is authored: the harness makes notes available, warns, and nudges, but does not validate whether citations entail claims or whether semantic repair and re-testing occurred. The level-transfer warning clearing after any edit is one bounded example, not evidence of campaign noncompliance. The replay relation is implemented and uses the heterogeneous checks listed in the ledger. HISTORY_FIT is bounded to those recorded-history checks; INCOMPLETE may support search only when the current modeled state remains usable; and MISMATCH categorically refuses search. Automatic replay therefore gives the executable representation a firmer episode/rule relation than note labels have, without licensing general explanatory warrant. (ARC skill doctrine; inspection and status; executable-rule replay and search.)

Plan freshness bounds use, not endorsement

ARC and Commonplace share an applicability function but use different transitions. ARC permits a plan only while its event, observation, and rules.py provenance is unchanged, then meters the queue through live results. Commonplace finalization pins review evidence to note and criterion snapshots, while acknowledgement can advance a baseline after a non-invalidating note change without changing the evidence pair. ARC has no corresponding acknowledgement transition, and Commonplace freshness does not execute actions. In neither system does applicability imply truth, endorsement, or handled findings. (Commonplace review-system semantics; live execution; executable-rule replay and search.)

Consequence-mediated explanation participation

ARC's explanation route is best described locally as consequence-mediated participation. Doctrine and status output make prose notes available to guide predictions and plans if the agent consumes them; the supplied evidence does not show realized campaign influence. Invoked executable rules have an implemented route into generated plans. In both cases, the hard live checks mostly evaluate observable consequences rather than explanation as explanation. A provenance-bound plan can therefore have strong operational containment while note status remains self-declared or replay remains incomplete.

The decisive selection loci in the selected implementation follow the same pattern. The harness rejects absent prediction admission, a contradictory executable model's access to search, a stale plan, or the remaining suffix of a diverging queue. It does not expose a code-enforced accept/reject transition over a population of prose theories. That matters under the definition of a proposal-selection loop: retention of predictions and surprises is not itself selection of accepted explanations. Consequence-mediated participation therefore names how explanation affects behavior in this case; it does not imply that consequence fit establishes explanatory-reach or that no informal theory comparison occurred.

Comparison with the five existing cases

This section compares ARC with the conclusions already recorded in the epistemic-architectures workshop framing, four-system baseline, AI Research OS reading, and operator correction. It does not independently re-audit the underlying Eigenius, ScienceFlow, AI Research OS, or ontology sources, and their evidence grades are not equal. In particular, the ontology draft had no observed running loop; the operator correction records its intended lab-tooling scope.

Existing case Operative selection/checking locus in the supplied comparison Explanation route and authority What ARC adds or changes
ScienceFlow A task-metric evaluator gates stage acceptance, and accepted anchors feed later iterations. The retained state lacks a claim or hypothesis object, so explanation is absent rather than weakly graded. ARC also has an environment outcome, but it adds notes designed and displayed for action guidance, executable models, and per-action consequence checks. ARC is not an absent-explanation case. (Four-system baseline; prediction parser and grader; live execution; executable-rule replay and search.)
Ontology draft A measurement policy operates over signed attestations; no running loop was observed. Hypothesis and Knowledge are represented but unscored and non-authoritative. The operator correction frames this as intended lab-tooling scope. ARC does not exile explanation. Notes are designed and exposed to guide behavior, and invoked executable models can generate plans, although hard live checks mostly land on predicted consequences rather than accepting explanations directly.
Eigenius Formal, type, and certificate checks create route-specific gaps. Mechanized faithfulness is capped below Verified. Explanatory content can enter later reasoning, but authority is bounded by epistemic grade. ARC also needs route analysis, but its distinctive containment is temporal and operational: plan provenance limits use, and an after-action surprise truncates the queue. It is not primarily an epistemic grade cap. (Four-system baseline; live execution; executable-rule replay and search.)
Commonplace Structural validation and verdict-kind gates can have acceptance force. Explanatory-reach critique is report-kind, and the manual reach audit is outside default acceptance. Explanation is the theoretical quality goal, but the comparison found it weakly operative. ARC's per-action consequence checks are more immediately operative. Its plan provenance resembles Commonplace review freshness only as an applicability boundary; neither freshness marker is endorsement. (Four-system baseline; explanatory-reach; Commonplace review-system semantics.)
AI Research OS Structural lint and read-time attention routing coexist with universal page retention and no reject-capable content acceptance. Explanation is the retained medium, marked but unscored and fully consumed, so its operational influence is uncontained by content acceptance. ARC also retains editable synthesis, but it preserves inspectable events, automatically replays executable rules, and halts queued actions after divergence. (AI Research OS reading; inspection and status; executable-rule replay and search; live execution.)
ARC as the sixth case Prediction presence controls admission; observation controls consequence grades and queue continuation; replay controls model-search eligibility; provenance controls plan applicability; the environment controls task completion. Predictions and invoked executable models have implemented routes; notes have designed and available authority whose realized campaign effect is not shown. These routes receive different checks. ARC puts several participation and containment profiles inside one system. Its contribution to the comparison is route-level and temporal asymmetry, not a ranking or a generic claim that ARC has “more rigorous prediction.”

The earlier five cases already put several meanings under “containment”: an epistemic grade cap, advisory force, explanation exclusion, or the absence of reject-capable acceptance. ARC adds several such profiles inside a single system. That makes the collapse untenable: the comparison cannot say merely that ARC is “contained” without naming what is contained, when, and with what force.

Commonplace-specific epistemic comparison

Commonplace supplies distinctions for interpreting ARC, not a claim to stronger empirical truth guarantees. ARC does not adopt the Commonplace discovery lifecycle, and HISTORY_FIT is not an assay outcome.

Commonplace distinction ARC route Comparison limit
The discovery lifecycle separates conjecture, consequence derivation, testing, acceptance, and integration. ARC implements consequence pre-registration and testing at action grain. Executable replay adds a named recorded-history test. The selected code has no acceptance transition for prose mechanics. HISTORY_FIT is a bounded test status, not general explanatory acceptance. (ARC skill doctrine; executable-rule replay and search; CLI routes.)
Fine-grained theory warrant keeps support at the claim, conjunction, model, and scope identified by evidence. One prediction grade bears on one parsed consequence and transition; replay bears on an integrated executable model under its comparison surface; the scorecard bears on the configured bundle. None automatically distributes warrant to prose theory, prediction gating, notes, planning, executable modeling, or the base model. (Prediction parser and grader; ARC repository README.)
Explanatory-reach requires more than familiar-case fit: load-bearing premises must remain criticizable and relevant rivals or transfers must be tested where the claim requires them. Next-frame predictions state criticizable consequences; executable rules expose a mechanism; live plan checks add prospective fit along encountered paths. Recorded-history replay does not by itself vary premises, eliminate relevant rivals, or test transfer outside observed and measured routes. Ordinary transitions may be compared only through observe, while terminal transitions use the model's goal or optional dead predicate. (Action-conditioned world-model tests; formal-system explanatory-reach.)
Commonplace review-system semantics make freshness a current snapshot-applicability record. Finalization pins new evidence; acknowledgement can advance an older evidence pair across a non-invalidating note change. Freshness is not endorsement or proof that findings were handled. ARC pins a plan to current event identity, observation, and rules, then checks each executed result. Both are applicability claims, but only ARC requires unchanged provenance and directly meters action continuation. That operational difference does not expand epistemic warrant. (Live execution; inspection and status.)

The useful exchange is asymmetric. Commonplace's vocabulary prevents ARC's operational containment from being mistaken for epistemic acceptance. ARC supplies a concrete external case in which that separation changes the comparison result.

Taxonomy gate: participation × containment 2×2 — FAIL unchanged

The system-level participation × containment 2×2 fails unchanged on the ARC case. ARC does not falsify participation or containment as questions. It falsifies the adequacy of assigning the whole system one unqualified cell.

The tested 2×2 comes from the AI Research OS reading and four-system baseline. The pressure test below applies the already-existing path distinctions in artifact-analysis axes, behavioral authority, action-model consumption paths, fine-grained theory warrant.

ARC route Explanation participation Epistemic containment Operational containment Why one system cell loses information
Prose notes Designed and displayed as action-guiding synthesis; realized use is unshown Self-declared status; no semantic gate in the selected routes Available to influence later actions if consumed, without a note-level accept/reject transition Assigned participation has no matching demonstrated epistemic containment.
Standalone prediction An observable forecast is required; an explanation may motivate it, but the optional reason is ungraded and no dependency is enforced Grade limited to one stated consequence on one transition Prediction presence gates admission, but a miss creates no persistent stop Pre-action admission and post-action grading have different force.
Manual batch A required observable forecast guides each step; its explanatory source is untracked The same transition-bounded grade First surprise truncates only the future suffix of the active queue Containment is temporal and queue-scoped; the surprising action is already spent.
Executable replay and search An executable action-conditioned transition model participates HISTORY_FIT is bounded by the applicable replay checks; explicit gaps remain INCOMPLETE MISMATCH categorically blocks search; INCOMPLETE proceeds only with a usable current model state Epistemic status and search eligibility do not coincide.
Generated plan A model-mediated action sequence participates Freshness does not endorse the plan's producing model Stale provenance refuses use; live divergence discards the suffix Operational containment can be strong without model acceptance.
Benchmark outcome Not applicable: the scorecard is an after-run evaluator, not explanatory content The outcome bears on the configured bundle and gives no theory or component attribution Environment completion ends the task The evaluated bundle is the target, not an explanation participating in this route.

“Contained” currently collapses at least epistemic status, applicability, search eligibility, and continuation horizon. A successor comparison must name the route, target, timing, and force, then state whether containment limits epistemic status, operational continuation, or both. Participation remains useful when attached to a route and qualified as direct or consequence-mediated. But the result does not require a permanent taxonomy cell for every route, invalidate every route-level 2×2, or warrant a universal replacement taxonomy from ARC alone.

Repository-reported campaign observations

The campaign figures bound this architecture reading; they do not establish component effects.

  • The ARC repository README reports 25/25 games, 183/183 levels, RHAE 100.00, and 7,645 actions against a reported median-human 17,135.
  • It reports 7,627 graded predictions, 443 misses, and at least one miss in every game. If accepted, these figures show that local prediction error and whole-system completion coexisted. They do not identify why completion occurred.
  • It reports that 91.6% of presses occurred in plans, with a 37.1% miss rate for single test presses and 2.9% for planned steps. The ARC skill doctrine routes exploration to single steps and “proven” mechanics to plans, so these are selected action classes, not an intervention on planning or halting.
  • It reports 115 context compactions and a median notes length of 60 lines. Those facts are compatible with notes serving as recovery state, but do not establish note quality, citation fidelity, doctrinal compliance, or causal necessity.
  • It reports offline Python use in 24/25 games. The shipped full rules route was attempted once and never fit, so the campaign supplies little positive worked evidence for that particular route despite its implemented contract.

The checkout supplied no run directories, so all quantities and campaign behavior above remain claims reported by the ARC repository README, not independently reconstructed event evidence. The score is a whole configured-system outcome. No matched comparison isolates prediction pre-registration, post-action halting, notes, executable modeling, planning, or base-model contribution. The test-versus-plan miss-rate contrast is confounded by doctrine-driven routing. It is therefore neither an ablation nor evidence that prediction gating, planning, or halting caused the reported result. (Experimental contrast limit; fine-grained theory warrant.)

Negative result: ARC does not test explanatory-quality underselection

The supplied evidence has no independently characterized candidate-theory population, no harness accept/reject boundary over prose explanations, and no calibrated explanation-quality outcome. The implemented selection force lands instead on action admission, model-search eligibility, plan applicability, and queue suffixes. Repository outcomes and prediction grades do not change that target. (Weak-oracle underselection conjecture; proposal-selection loop; ARC repository README; prediction parser and grader; live execution; executable-rule replay and search.)

ARC should therefore not become positive evidenced-by support for the explanatory-quality underselection conjecture on this record. This is an evidential limitation, not a refutation of the conjecture, a judgment about ARC's explanations, or proof that no informal theory comparison occurred. Run artifacts could independently verify campaign events or compliance, but a component-benefit claim would still require a comparison that varies the relevant component.

Bottom line and handoff

ARC is useful as the sixth case because it shows why a system-level label transfers results among unlike routes. Prediction presence gates admission; observations grade consequences; some surprises truncate a queue; replay controls model-search access; provenance controls plan use; and the environment controls task completion. These results do not share one target, timing, or force. The comparison should keep participation and containment only as route-qualified questions and should report epistemic authority separately from operational authority.

No new participation × containment note is warranted. The separately authorized fold was completed on 2026-08-31: the behavioral-authority decomposition proposal now carries ARC as a bounded worked case grounded in the pinned live-action implementation ingest, not in this temporary workshop. The case shows that one real route family exercises the proposal's existing distinctions; it does not settle the proposal's open design choices or validate a universal decomposition.