oh-my-pi agentic-system analysis

Type: types/agentic-system-analysis-result.md

Run identity

Run state: kb/reports/state/agentic-system-analysis/AAS-2026-09-26-oh-my-pi-01/run-state.md

Generated review: kb/agentic-systems/reviews/oh-my-pi.md

Memory analysis report: kb/reports/state/agentic-system-analysis/AAS-2026-09-26-oh-my-pi-01/memory-report.md

Memory analysis report SHA-256: ab3c0949d73a749978fffd85bbef423b1f444c9ff3baef0071660b1fc9d7259c

Run AAS-2026-09-26-oh-my-pi-01 uses only frozen primary sources and the required fresh memory specialist. The public review will reference the identical retained result under kb/reports/retained/agentic-system-analysis/AAS-2026-09-26-oh-my-pi-01/result.md. These are intended outputs; completion is declared by the run state. No prior-review prose or audit findings informed construction. Analysis model: runtime identity GPT-6; specialist identity is recorded in its typed report.

Boundary and evidence

Evidence basis: static implementation and shipped doctrine at Git commit be6cb8217cd4c1dafcc86793ae5d809ea4d7396a, analysed 2026-09-26. This locally available commit is dated 2026-09-05; package metadata reports version 18.1.10. A later fetch completed after the boundary was frozen; fetched revisions were not substituted or analysed. This result does not assert current-tip coverage.

The intended use is understanding what oh-my-pi wires as an enclosing coding-agent runtime: entry modes, scheduling, context, tools, delegation, persistence, and material product/instruction revision mechanisms. It includes the shipped CLI and SDK, agent core, local context/memory code and client-facing backend contracts. It includes the operator where an inspected route requires a human proposal, veto or adoption decision, and names that contribution explicitly.

The boundary is whole-system at the inspectable runtime level, not exhaustive verification of every tool. Excluded provider internals prevent conclusions about actual parameter fixity, private reasoning and provider-side adaptation. Excluded external memory-service implementations prevent reconstruction of their opaque representation and admission algorithms. Uninspected external MCP servers, language servers, OS controls, browser/debugger deployments and host clients prevent guarantees about their execution semantics and containment. Arbitrary user extensions are represented by their host interface, not audited individually. No user configuration, credentials, execution traces or measured outcomes were inspected; wiring is the strongest operational evidence tier here. Benchmark blog claims and videos outside the repository are excluded.

Source register

Source ID Kind Identity/location Revision Evidence layer Inspected scope Citation anchors Access gaps and conclusion prevented
SRC-1 Git https://github.com/can1357/oh-my-pi be6cb8217cd4c1dafcc86793ae5d809ea4d7396a Implementation Selected runtime, SDK, extension/tool policy, task, session, memory, compaction and revision files cited by canonical records Full commit-relative paths on each record; quote blocks resolve complete pinned blobs No deployed run, provider internals or external service internals; no observed activation, causal benefit or global isolation finding
SRC-2 Git https://github.com/can1357/oh-my-pi be6cb8217cd4c1dafcc86793ae5d809ea4d7396a Doctrine/design, with README performance assertions kept as attributed claims README.md, package metadata, shipped prompts and documentation cited by records Same pinned repository; full paths on records Intent and advertised results do not establish operation or comparative performance

Operational access root: /home/zby/llm/commonplace/related-systems/can1357--oh-my-pi. Source allowlist: this repository/revision only. Git inspection used full-commit show, grep and ls-tree with replacement objects disabled; source worktree contents and current HEAD were not evidence. Metadata anchor: SRC-2 packages/coding-agent/package.json:1-32. No probe source records exist.

Shared records

Components

CMP-1 — Session/agent driver. Symbolic TypeScript machinery connects CLI modes, createAgentSession, AgentSession and the lower-level agent loop. It owns scheduling, context assembly and tool dispatch, rather than the model provider. Conclusion status: wired. SRC-1 packages/coding-agent/src/main.ts:2045-2118, packages/coding-agent/src/sdk.ts:3531-3645, packages/agent/src/agent-loop.ts:1380-1565.

CMP-2 — Configured chat models, also used by delegated workers and model-backed auxiliary calls. Distributed-parametric computation lies behind provider/model identifiers, a model registry and credential resolver. Inference dispatch: wired. Exact immutable parameter/version pinning: uninspected; a full source commit pins the adapter, not the remote model weights. Provider-side parameter changes during operation: uninspected. The inspected SDK passes prompts, tools, sampling options and credentials; it does not expose the provider's training process. Do not infer learning or its absence from that interface. SRC-1 packages/coding-agent/src/sdk.ts:3531-3593, packages/coding-agent/src/task/executor.ts:3330-3352.

getApiKey: options.getApiKey ?? (requestModel => modelRegistry.resolver(requestModel, agent.sessionId)), --- packages/coding-agent/src/sdk.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

CMP-3 — Tool registry and effect executors. Built-in, extension, SDK and provider bridge tools participate in dispatch; shell commands, files, language servers, external tools and host clients supply effects outside the model. Symbolic mechanisms; conclusion status: wired. Current deployment grants and OS isolation: uninspected. SRC-1 packages/coding-agent/src/sdk.ts:2918-2930, packages/agent/src/agent-loop.ts:2650-2697.

CMP-4 — Extension runtime. Host-loaded factories register tools and hooks, inject messages and expose command execution. Symbolic code with runtime-configured content. Conclusion status: wired. The extension author and host deployment are separate trust actors; no claim of containment follows from an intercepted tool call. SRC-1 packages/coding-agent/src/extensibility/extensions/loader.ts:258-290,361-385.

CMP-5 —Local rollout extractor/consolidator and continuation summarizers. LLM inference receives traces and writes text/artifacts, not weights. Local phases call completeSimple with runtime-resolved models; branch summarizer uses current model. Exact deployment model/version and endpoints are configuration-dependent, not pinned by this source revision. No parameter-changing call was found in these inspected routes; external provider internals remain unknown. Source: SRC-1 packages/coding-agent/src/memories/index.ts:724-769,872-910; packages/coding-agent/src/session/agent-session.ts:9471-9496. SRC-2 compaction prompt above names continuation purpose. Consolidation is not itself evidence of model training.

CMP-6 — Mnemopi extractor and embedding models. Configured host/remote completion extracts fact categories; optional ONNX/remote embedding inference produces access vectors. Wrapper defaults differ from the package README: actual wrapper selects BAAI/bge-base-en-v1.5 or intfloat/multilingual-e5-large, with explicit/environment override. These are model identifiers, not immutable weight digests. Ordinary sleep consolidation explicitly marks llm_used false. There is no inspected weight-update operation; advanced inference algorithms are partly uninspected. Source: SRC-1 packages/coding-agent/src/mnemopi/config.ts:43-103; packages/coding-agent/src/mnemopi/state.ts:32-65; packages/mnemopi/src/core/extraction.ts:373-410; packages/mnemopi/src/core/beam/consolidate.ts:1020-1059.

const variantModel = embeddingVariant === "multilingual" ? "intfloat/multilingual-e5-large" : "BAAI/bge-base-en-v1.5"; // Precedence: explicit mnemopi.embeddingModel setting > MNEMOPI_EMBEDDING_MODEL // env (documented model-level override) > variant-derived default. Without the env // term a variant default would silently shadow a user's configured env model. const embeddingModel = embeddingOverride?.trim() || Bun.env.MNEMOPI_EMBEDDING_MODEL?.trim() || variantModel; --- packages/coding-agent/src/mnemopi/config.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

if (configuredLlmWillHandleCall()) { diag.recordAttempt("host"); try { const raw = await callConfiguredCompletion(prompt, 0, { maxTokens: llmMaxTokens(), task: { kind: "memory-extraction", input: text }, }); --- packages/mnemopi/src/core/extraction.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

CMP-7 — Sharpshooter extractor/consolidator. Configured model, falling back to smol role, performs two instruction-governed inference calls; queue and file replacement are the mutable products. No weight-update route in inspected files, and no fixed deployed endpoint/model revision supplied. Source: SRC-1 packages/coding-agent/src/sharpshooter/extract.ts:151-168; packages/coding-agent/src/sharpshooter/consolidate.ts:139-150.

CMP-8 — Hindsight remote retain/recall/reflect and provider-native compaction. Host sends API inputs and consumes results; upstream model identity, parameter updates, storage and algorithm are unknown. Compaction payload can preserve encrypted reasoning. Reflect has an important backend distinction: Hindsight returns service synthesis, while Mnemopi's tool formats recalled context without a dedicated reflection-model call. Service availability is configurable, not evidence of a fixed version. Source: SRC-1 packages/coding-agent/src/hindsight/client.ts:36-68; packages/coding-agent/src/tools/memory-reflect.ts:33-80; packages/agent/src/compaction/openai.ts:1-15; packages/coding-agent/src/session/session-context.ts:165-177.

              const summary = state.formatContextScoped(results);
              return {
                  content: [{ type: "text", text: `Based on recalled memories:\n\n${summary}` }],
                  details: {},
              };

--- packages/coding-agent/src/tools/memory-reflect.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Operative objects

OBJ-1 — External project files and task patches. Symbolic code and potentially natural-language documents on files/repositories; work products rather than automatically the harness's own machinery. Product revision is covered by the editing and task routes. SRC-1 packages/coding-agent/src/task/isolation-runner.ts:350-461.

OBJ-2 — Current tool calls and returned results. Structured call identity/arguments plus textual or other result content; current invocation state. A retained session containing these has a separate memory role. Source truth, model explanation and tool exit status remain distinct. SRC-1 packages/agent/src/agent-loop.ts:1380-1453,2650-2697.

OBJ-3 — Active tool registry, settings and extension factories. Symbolic operational configuration and executable capabilities, not truth-apt claims merely because they are retained. Changes affect admission, exposure and execution; default availability is distinct from the current deployment's grants. SRC-1 packages/coding-agent/src/sdk.ts:2795-2815,2918-2930, packages/coding-agent/src/extensibility/extensions/loader.ts:277-290,361-385.

ID Source-native object and memory classification Evidence
OBJ-12 Session entries: messages/tool results, branch parents, leaf IDs and checkpoints. Raw traces coexist with derived entries; JSONL files by default, in-memory option, Redis or SQL SDK backends. Entry/parent IDs route later context; content is conversation evidence. SRC-1 packages/coding-agent/src/session/session-storage.ts:15-91; packages/coding-agent/src/session/session-context.ts:195-236; SRC-2 packages/coding-agent/examples/sdk/12-redis-sessions.ts:1-54, packages/coding-agent/examples/sdk/13-sql-sessions.ts:1-61.
OBJ-13 Derived checkpoint family: soft/handoff/branch prose, remote replacement history with opaque encrypted content, snapcompact PNG/text archive, and shake placeholders with recoverable artifacts. Retained in session entries/artifact files. Prose is natural language, IDs/schema symbolic, raster language only partially maps; remote payload form unknown. Consumer is resumed model/provider conversion. SRC-1 packages/coding-agent/src/session/session-manager.ts:2409-2438; packages/coding-agent/src/session/session-maintenance.ts:650-709,1492-1535; packages/coding-agent/src/session/session-context.ts:424-475; packages/snapcompact/src/snapcompact.ts:1819-1867,2037-2056,2100-2148.
OBJ-14 Local rollout memory: SQLite stage-one outputs/job watermarks; project-root raw_memories.md, rollout summaries, MEMORY.md, memory_summary.md and generated skill files; separate learned.md lessons. Raw-memory label denotes extracted material, not original logs. Summary and lessons are selected for prompt injection; other files require agent reads. Generated skill presence alone is not executable adoption. SRC-1 packages/coding-agent/src/memories/index.ts:123-228,345-383,477-512,724-810,872-1009,1280-1297,1338-1364; SRC-2 packages/coding-agent/src/prompts/memories/read-path.md:1-17.
OBJ-15 Mnemopi working/episodic rows and extracted fact categories, scoped SQLite banks, FTS/vector access structures, source/session IDs, summary_of lineage, veracity/importance/expiry and graph ingestion. Natural-language content plus symbolic metadata; embeddings are retrieval representations, not evidence of trained model updates. SRC-1 packages/coding-agent/src/mnemopi/state.ts:324-369,472-551; packages/mnemopi/src/core/beam/consolidate.ts:389-432,524-550,972-1059; packages/mnemopi/src/core/beam/recall.ts:731-774,963-1013; packages/coding-agent/src/mnemopi/config.ts:43-103.
OBJ-16 Hindsight bank documents, recall results and operator-curated mental models, exposed as service objects; local in-memory snippets/cursors. Readable API text is visible, upstream storage/derivation algorithms are excluded. SRC-1 packages/coding-agent/src/hindsight/state.ts:293-373,425-523; packages/coding-agent/src/hindsight/client.ts:36-68; packages/coding-agent/src/hindsight/backend.ts:86-123.
OBJ-17 Sharpshooter delta queue retains statement/evidence/friction flags and optional rationale/rejected alternative; three project decision files architecture.md, product.md, style.md contain consolidated instructions. Queue provenance and final normative prose are distinct objects within this family. SRC-1 packages/coding-agent/src/sharpshooter/extract.ts:29-46,98-148,256-293; packages/coding-agent/src/sharpshooter/consolidate.ts:56-78,123-150; SRC-2 packages/coding-agent/src/prompts/memories/sharpshooter-consolidate-system.md:1-43.

Memory payload and authority evidence

The runtime has two different continuity mechanisms: retained session history is rebuilt around the selected branch and latest checkpoint, while memory backends supply project/bank material across conversations. Five ordered context-maintenance choices include provider compaction, raster archives, prose handoff, mechanical shrinking and prose summarization. Their outputs cannot all be called text summaries.

export const DEFAULT_COMPACTION_METHOD_ORDER: CompactionMethod[] = [ "remote", "snapcompact", "handoff", "shake", "soft", ]; --- packages/coding-agent/src/session/compaction-methods.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

A checkpoint's readable summary is only one part of the object. Provider replacement history can carry encrypted content. Snapcompact preserves text as PNG frames plus readable edges, and the later vision model receives those images again. Classification must include the operative payload, not only what the interface displays.

const candidate = compaction?.preserveData?.openaiRemoteCompaction; if (!isRecord(candidate)) return undefined; if (typeof candidate.provider !== "string" || candidate.provider.length === 0) return undefined; if (!Array.isArray(candidate.replacementHistory) || !candidate.replacementHistory.every(isRecord)) return undefined; return { type: "openaiResponsesHistory", provider: candidate.provider, items: candidate.replacementHistory, }; --- packages/coding-agent/src/session/session-context.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

    • OpenAI remote compaction V1 (/responses/compact): preserves encrypted
  • reasoning across compactions by submitting the full responses-API native
  • history and storing the returned compaction / compaction_summary
  • item in preserveData so future turns can replay the encrypted state. --- packages/agent/src/compaction/openai.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

export function images(archive: Archive): ImageContent[] { return archive.frames.map(frame => ({ type: "image", data: frame.data, mimeType: frame.mimeType, ...(frame.detail ? { detail: frame.detail } : {}), })); --- packages/snapcompact/src/snapcompact.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Local project memory and sharpshooter decisions have deliberately different authority. Local memory asks the agent to verify against the repository. Sharpshooter puts selected project decisions into the instruction channel and says to follow them unless the user overrides. Neither instruction demonstrates compliant behavior.

  1. Memory: heuristics/process context; current repo files, runtime output, user instruction: factual state/final decisions.
  2. Memory changes plan → cite artifact path (e.g. memory://root/skills/<name>/SKILL.md) and current-repo evidence.
  3. Memory disagreement with repo state/user instruction → stale; corrected behavior, then update/regenerate memory artifacts.
  4. Confidence only after repository verification; memory alone NEVER sufficient proof. --- packages/coding-agent/src/prompts/memories/read-path.md @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

    const parts = [ "Project decision memory (sharpshooter). These are friction-earned decisions; follow them unless the user overrides.", ]; for (const file of populated) parts.push(## ${file.name.slice(0, -3)}\n\n${file.content.trim()}); return truncateApproxTokens(parts.join("\n\n"), settings.get("sharpshooter.injectionTokenLimit")); --- packages/coding-agent/src/sharpshooter/backend.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Memory storage and authoring evidence

The storage implementation for OBJ-12 supports files through the file-backed session implementation and in-memory through a path-indexed Map. Redis stores serialized content by key; SQL stores content in a table. SDK examples cited on OBJ-12 wire those adapters to named create/resume agent consumers; their weaker afforded basis governs the union.

export class FileSessionStorage implements SessionStorage { ensureDirSync(dir: string): void { if (!fs.existsSync(dir)) { fs.mkdirSync(dir, { recursive: true }); } --- packages/coding-agent/src/session/session-storage.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

export class MemorySessionStorage implements SessionStorage { // Each path keeps appended string chunks plus cumulative UTF-8 byte offsets. // Full reads materialize the chunks into one string chunk, so repeated reads // do not re-join stale history. Later appends still stay O(1) by pushing // after that materialized chunk. Prefix/suffix reads binary-search byte // offsets and join only the requested window. #files = new Map(); --- packages/coding-agent/src/session/session-storage.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

const WRITE_FULL_SCRIPT = -- OMP_WRITE_FULL redis.call("SET", KEYS[1], ARGV[1]) redis.call("HSET", KEYS[2], ARGV[2], ARGV[3]) ---packages/coding-agent/src/session/redis-session-storage.ts@be6cb8217cd4c1dafcc86793ae5d809ea4d7396a`

export type SqlSessionStorageAdapter = "postgres" | "mysql" | "sqlite"; --- packages/coding-agent/src/session/sql-session-storage.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

  createTable:
      `CREATE TABLE IF NOT EXISTS ${table} (` +
      `path TEXT PRIMARY KEY, ` +
      `content TEXT NOT NULL, ` +
      `mtime_ms ${mtimeType} NOT NULL, ` +

--- packages/coding-agent/src/session/sql-session-storage.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

For storage, OBJ-14 additionally establishes files through the retained MEMORY.md/memory_summary.md write quote above and SQLite through the actual database import/construction and schema. OBJ-15 adds persisted graph edges and vectors as access structures; these do not imply that the model parameters are trained.

import type { Database } from "bun:sqlite"; --- packages/coding-agent/src/memories/index.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

export function openMemoryDb(dbPath: string): Database { const db = new Database(dbPath); // Install the busy handler BEFORE any lock-taking statement. See #2421. db.exec("PRAGMA busy_timeout = 5000"); --- packages/coding-agent/src/memories/storage.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

CREATE TABLE IF NOT EXISTS stage1_outputs ( thread_id TEXT PRIMARY KEY, source_updated_at INTEGER NOT NULL, raw_memory TEXT NOT NULL, rollout_summary TEXT NOT NULL, --- packages/coding-agent/src/memories/storage.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

  this.db.run(
      `INSERT INTO graph_edges (source, target, edge_type, weight, timestamp)
          VALUES (?, ?, ?, ?, ?)
          ON CONFLICT(source, target, edge_type) DO UPDATE SET
              weight = excluded.weight,
              timestamp = excluded.timestamp`,

--- packages/mnemopi/src/core/episodic-graph.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

  using insertEmbedding = beam.db.prepare(
      "INSERT OR REPLACE INTO memory_embeddings(memory_id, embedding_json, model) VALUES (?, ?, ?)",
  );

--- packages/mnemopi/src/core/beam/helpers.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

At the API boundary, OBJ-16 is a service-object classification at the inspected API boundary: a bank-addressed request persists memory content. It is not an inference about the server's database substrate.

  return this.#request<RetainResponse>(
      "POST",
      `/v1/default/banks/${encodeURIComponent(bankId)}/memories`,
      "retain",
      {
          body: { items: [item], async: options?.async },
          signal: options?.signal,
          timeoutMs: this.#retainTimeoutMs,

--- packages/coding-agent/src/hindsight/client.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

The manual-authoring branch of RTE-17 affords manual authoring through editable lesson files. Read-time normalization explicitly accounts for hand edits and the previously quoted injection route consumes their text. This is manual agency; it is separate from human-triggered automatic extraction.

/* * Read learned.md, neutralizing each line on read too — a hand-edited or * pre-existing file bypasses write-time normalization and the block renders * unescaped into the system prompt. Returns "" when absent/unreadable. / --- packages/coding-agent/src/memories/index.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

ID Source-native object and lineage Form and truth-apt content Role, evidence and limit
OBJ-4 FileDiagnosticsResult returned through diagnosticsJson for that edit; LSP/linter reports consumed in tool output Structured observations/claims about a particular file under external checker semantics Product feedback; SRC-1 packages/coding-agent/src/lsp/diagnostics.ts, waitForDiagnostics/getDiagnosticsForFile and packages/coding-agent/src/edit/index.ts; no language-server truth guarantee
OBJ-5 Reviewer incremental findings and overall_correctness/explanation/confidence for an assigned diff; model produces, parent/operator consumes Natural-language bug claims and verdict, structured fields; ampliative correctness/impact judgments Patch criticism; SRC-2 packages/coding-agent/src/prompts/agents/reviewer.md; SRC-1 packages/coding-agent/src/tools/review.ts, parseFindingDetails; structural validity does not prove a bug
OBJ-6 Advisor note plus nit/concern/blocker severity; advisor produces, primary receives advisory/steering Natural-language criticism, potentially truth-apt but arbitrary note content makes transformation indeterminate Ongoing feedback; SRC-1 packages/coding-agent/src/advisor/advise-tool.ts and packages/coding-agent/src/session/session-advisors.ts, routeAdvice; delivery acceptance is emission-budget admission
OBJ-7 Generated TTSR markdown with condition/scope/body; complaint and assistant history feed ephemeral model; TtsrManager and primary consume saved rule Mixed symbolic matcher and prescriptive guidance; no required truth-apt efficacy proposition Correct recurring behavior; SRC-1 packages/coding-agent/src/modes/controllers/omfg-controller.ts and packages/coding-agent/src/modes/controllers/omfg-rule.ts; SRC-2 packages/coding-agent/src/prompts/system/omfg-user.md
OBJ-8 Autoresearch changed program variant and described experimental change on active branch; primary generates, fixed harness evaluates, later iterations inherit kept version Formal program plus candidate conjecture that modification improves target metric while preserving correctness Optimization candidate; SRC-2 packages/coding-agent/src/autoresearch/prompt.md; SRC-1 packages/coding-agent/src/autoresearch/tools/log-experiment.ts; description need not state an explanatory hypothesis
OBJ-9 Autoresearch run record: captured output, parsedPrimary, supplied metric, status, commitHash, modifiedPaths, flags and confidence Structured measurements and recorded agent judgments; captured and supplied metrics have different producers and remain distinguishable Evidence and keep/discard record; SRC-1 packages/coding-agent/src/autoresearch/tools/run-experiment.ts and packages/coding-agent/src/autoresearch/tools/log-experiment.ts
OBJ-10 Autoresearch session.notes playbook, hypotheses and ideas; primary supplies body/append_idea; later system prompt consumes it Mixed prose; may reshape observations, conjecture, or prescribe policy; indeterminate without instance Durable experiment guidance; SRC-1 packages/coding-agent/src/autoresearch/tools/update-notes.ts and packages/coding-agent/src/autoresearch/index.ts, before_agent_start

OBJ-11 — Managed SKILL.md procedures: natural-language instructions plus symbolic frontmatter on files, with isolated managed directory and lower precedence than authored skills. Producer is a model/tool call, optionally after an auto-learn capture turn; consumer is future skill discovery and the model that loads/uses a selected skill. Excluded from the bounded backend comparison profile. SRC-1 packages/coding-agent/src/autolearn/managed-skills.ts:152-212, packages/coding-agent/src/extensibility/skills.ts:376-410; SRC-2 packages/coding-agent/src/prompts/system/autolearn-guidance.md.

Storage supplements: OBJ-4 and OBJ-6 are current-run reports/messages with session retention where used; OBJ-5 is structured reviewer output; OBJ-7 is project/global rule-file content; OBJ-8 is repository code, and OBJ-9/OBJ-10 are autoresearch retained records/notes. These distinctions do not assert epistemic endorsement. Source anchors remain on their routes.

Routes

RTE-1 — Ordinary prompt → provider → tools → next turn → final stream. Conclusion status: wired. Trigger/principal: operator prompt or client request. Identity: session/provider session IDs and per-call tool IDs. Next-step owner: Agent loop; decision policy combines model-generated calls with symbolic scheduling, stop reasons, required-tool rules, deadline and queues. Context is transformed, converted and normalized before provider dispatch. The executor receives actual arguments and call-scoped context; returned results re-enter the conversation. Immediate return: streamed progress/results and final messages. Later read-back: the session routes described by the memory records, not the ephemeral loop alone. Delegated visibility: explicit task context through RTE-3, not automatic visibility of every parent message. Selection predicate: runnable stop plus pending calls, then steering/asides/follow-ups. Expiry: per-call deadlines and abort signals; persistent history has separate retention. Effect: tool execution is wired; material effect and model activation are uninspected without a run. Recovery: errors become tool results; aborted/length-truncated calls receive paired placeholders; external abort leaves steering queued. Terminal output follows empty queues or deadline/abort. SRC-1 packages/agent/src/agent-loop.ts:1380-1565,1593-1663,2650-2697.

if (config.transformContext) { messages = await config.transformContext(messages, signal); }

const llmMessages = await config.convertToLlm(messages); --- packages/agent/src/agent-loop.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

const runnableStop = message.stopReason === "toolUse" || message.stopReason === "stop"; hasMoreToolCalls = runnableStop && toolCalls.length > 0; --- packages/agent/src/agent-loop.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

const followUpMessages = signal?.aborted ? [] : (await config.getFollowUpMessages?.(signal)) || []; if (lateSteering.length > 0 || asideMessages.length > 0 || followUpMessages.length > 0) { --- packages/agent/src/agent-loop.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-2 — Intercepted tool-call admission. Conclusion status: wired. Trigger: dispatch of a registered tool. Next-step owner: ExtensionToolWrapper plus configured hook handlers; policy form: symbolic tier/override resolution with optional human UI judgment. Original explicit denial short-circuits; hook-revised arguments face another policy resolution. Proposed change: tool effect, which may modify OBJ-1 or state. Admission/rejection: deny, prompt or allow; user can refuse a required prompt, hook can block, and required approval without UI throws. Default wrapper fallback is yolo, not a deployed grant assertion. Persistence: settings remain policy inputs and results may enter session history; approval is call-local. Immediate return: effect or error. Later read-back: error/result through RTE-1 and session history. Delegated visibility: workers use RTE-3's different mode. Selection: tool identity/tier, effective args and override; expiry: call completion. Activation/effect is wiring only. Recovery: rejection precedes effect; no universal rollback of an already-executed command is established. Guidance is operational policy, not an acceptance criterion for the truth of an answer. Guarantee owner/enforcement: wrapper, protocol strength, registered dispatch paths; external contract: honest tool declarations and UI/provider behavior. Direct extension execution and shell-internal effects are outside this call-level proof. SRC-1 packages/coding-agent/src/sdk.ts:2795-2815,2918-2930, packages/coding-agent/src/extensibility/extensions/wrapper.ts:195-340, packages/coding-agent/src/tools/approval.ts:125-173.

for (const tool of toolRegistry.values()) { toolRegistry.set(tool.name, new ExtensionToolWrapper(tool, extensionRunner)); } --- packages/coding-agent/src/sdk.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

const resolvedArgs = approvalArgs(effectiveParams, context); const resolved = resolveApproval(this.tool, resolvedArgs, approvalMode, userPolicies); context?.xdevTierResolved?.(resolved.tier); if (resolved.policy === "deny") { throw denyError(resolved, this.tool.name); } --- packages/coding-agent/src/extensibility/extensions/wrapper.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

if (!this.runner.hasUI()) { const reason = "no interactive UI available"; await emitApprovalResolved(false, reason); --- packages/coding-agent/src/extensibility/extensions/wrapper.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-3 — Delegated task sessions and optional isolated apply-back. Conclusion status: wired. Trigger: task or eval-agent spawn; next-step owner: executor and child AgentSession, with child model deciding its tool sequence. Spawn policy resolves permitted agent identities; children receive selected tools, role prompt, supplied context, plan reference, model configuration and inherited credentials. Headless workers override approval mode to yolo; explicit user tool policies remain effective. The source's read-only classifier is an allowlist containing memory-writing tools, so its label does not mean no durable state mutation. Isolation selects a separate working directory/baseline; it is not evidence of an OS security boundary. Proposed changes: child product edits. Admission: optional merge/apply-back after run status and Git applicability checks; failed branches/patches remain for recovery. Those checks establish operational application, not task correctness. Return: child output, artifacts and apply-back result. Persistence: output/session and branch/patch artifacts; non-isolated worker sessions can be revived, isolated worktrees are torn down. Later read-back: resumed worker state and parent result consumption; delegated visibility is explicit supplied context and discovered sources. Selection: agent/tool/model policy and caller isolation options. Expiry: deadlines, termination and lifecycle cleanup; durable artifacts have separate retention. Guidance: task instructions and any project rules; localized proposed solutions are afforded, content-directed criticism and improved capacity are uninspected at this generic route. No supplied answer oracle is established for ordinary open tasks. SRC-1 packages/coding-agent/src/task/executor.ts:941-985,3323-3375,3429-3467, packages/coding-agent/src/task/spawn-policy.ts:18-62, packages/coding-agent/src/task/read-only-policy.ts:1-29, packages/coding-agent/src/task/isolation-runner.ts:350-461.

"tools.approvalMode": "yolo", --- packages/coding-agent/src/task/executor.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

requireYieldTool: true, contextFiles: options.contextFiles, skills: options.skills, promptTemplates: options.promptTemplates, workspaceTree: options.workspaceTree, rules: options.rules, --- packages/coding-agent/src/task/executor.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

const succeeded = result.exitCode === 0 && !result.error && !result.aborted; --- packages/coding-agent/src/task/isolation-runner.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

} else if (forwardApplies) { changesApplied = true; try { await repo.applyPatch(normalized, {}); hadAnyChanges = true; --- packages/coding-agent/src/task/isolation-runner.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-4 — Extension capability and context changes. Conclusion status: wired. Trigger: host discovers/imports an extension or runs a factory; owner: host module loader and extension author. Proposal: executable hooks, tools, provider registrations or injected messages. Admission: loading/running configured code; factory failure restores queued provider registrations. Rejection: load/factory failure or operator removal; there is no inspected semantic test for an extension's correctness. Recovery is narrow: queue rollback does not establish undo of arbitrary factory side effects. Immediate return: runtime registration/command result; later read-back: loaded configuration on future startup and injected messages on later turns. Delegated visibility: child session options may carry prepared extension paths, with restricted-tool branches excluding them. Selector: configured discovery paths and runtime calls; expiry: reload/disposal or removal, not content truth. Effects include direct command execution, additional capabilities and active-tool changes. Guidance is author/operator code and policy; theory-directed proposal/criticism, answer oracle and improvement attribution are uninspected. Loading an extension is capability revision, not evidence that the harness autonomously improved itself. SRC-1 packages/coding-agent/src/extensibility/extensions/loader.ts:258-290,361-385, packages/coding-agent/src/task/executor.ts:3352-3364.

exec(command: string, args: string[], options?: ExecOptions) { return execCommand(command, args, options?.cwd ?? this.cwd, options); } --- packages/coding-agent/src/extensibility/extensions/loader.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

try { await factory(api); } catch (error) { runtime.pendingProviderRegistrations.splice( 0, runtime.pendingProviderRegistrations.length, ...providerRegistrationCheckpoint, ); throw error; } --- packages/coding-agent/src/extensibility/extensions/loader.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-15 —Session continuity; wired. Producer appends message/checkpoint entries. Session load, resume or branch navigation selects a leaf, walks parent identities and chooses retained messages around the latest checkpoint. Main agent receives rebuilt context automatically; operator selection supplies a routing input, not a model pull request. Branch navigation can call a model to summarize the abandoned branch, attach it at the destination, and replace current messages. Source: SRC-1 packages/coding-agent/src/session/session-context.ts:195-236,424-495; packages/coding-agent/src/session/agent-session.ts:9471-9496,9569-9602. SQL/Redis have explicit SDK agent consumers but no supplied deployment: those alternatives are afforded.

const seenPathIds = new Set(); let current: SessionEntry | undefined = leaf; while (current && !seenPathIds.has(current.id)) { seenPathIds.add(current.id); path.push(current); current = current.parentId ? byId.get(current.parentId) : undefined; } path.reverse(); --- packages/coding-agent/src/session/session-context.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

  // Update agent state — build display context to populate agent messages.
  const stateContext = this.sessionManager.buildSessionContext();
  const displayContext = deobfuscateSessionContext(stateContext, this.#obfuscator);
  this.agent.replaceMessages(displayContext.messages);

--- packages/coding-agent/src/session/agent-session.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-16 — Trace compaction and continuation; wired. Manual/automatic context maintenance produces OBJ-13 from OBJ-12 and records it through appendCompaction; later rebuild delivers the retained replacement and kept tail to the model. Soft/handoff prompts preserve goals, decisions and rationale. Branch summaries use selected departed entries. Snapcompact mechanically renders conversation text; shake saves heavy regions as artifacts, replaces them with recovery references and rewrites entries. All are trace-fed transformations with later consumers; opaque provider payloads prevent a complete distilled-form judgment. Source: SRC-1 packages/coding-agent/src/session/session-maintenance.ts:650-709,1492-1535; packages/coding-agent/src/session/session-context.ts:424-475; SRC-2 packages/agent/src/compaction/prompts/compaction-summary.md:1-38, packages/agent/src/compaction/prompts/handoff-document.md:1-45.

  applyShakeRegions(items);
  this.#host.recordAnchoredHistoryRewrite(anchoredTokensRemoved);

  await this.#host.sessionManager.rewriteEntries();
  const sessionContext = this.#host.buildDisplaySessionContext();
  this.#host.agent.replaceMessages(sessionContext.messages);
  this.#host.resetAdvisorRuntimes("shake");
  this.#host.closeCodexProviderSessionsForHistoryRewrite();

--- packages/coding-agent/src/session/session-maintenance.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Key Decisions

  • [Decision]: [Brief rationale] --- packages/agent/src/compaction/prompts/compaction-summary.md @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-17 — Local two-stage rollout memory and lesson capture; wired. Startup of enabled top-level persistent sessions scans idle previous session files, excluding current thread; bounded model extraction writes stage-one memory/summary, then another model processes selected project outputs into replacement files. Main agent automatically receives bounded summary plus lessons on prompt rebuild; it can request detailed files through memory://. Files remain hand-editable; learned.md explicitly sanitizes hand-edited input on read. Source: SRC-1 packages/coding-agent/src/memories/index.ts:123-150,182-228,277-293,345-383,477-512,643-689,724-807,872-1009,1406-1419; packages/coding-agent/src/sdk.ts:3098-3106; SRC-2 packages/coding-agent/src/prompts/memories/read-path.md:1-17.

  const rawMemory = redactSecrets(schemaOutput.raw_memory).trim();
  const rolloutSummary = redactSecrets(schemaOutput.rollout_summary).trim();
  const rolloutSlug = schemaOutput.rollout_slug === null ? null : redactSecrets(schemaOutput.rollout_slug).trim();
  if (!rawMemory || !rolloutSummary) {
      return { kind: "no_output" };

--- packages/coding-agent/src/memories/index.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

await Bun.write(path.join(memoryRoot, "MEMORY.md"), ${consolidated.memoryMd.trim()}\n); await Bun.write(path.join(memoryRoot, "memory_summary.md"), ${consolidated.memorySummary.trim()}\n); const skillsDir = path.join(memoryRoot, "skills"); await fs.mkdir(skillsDir, { recursive: true }); --- packages/coding-agent/src/memories/index.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

const learnedBudget = Math.max(0, cfg.summaryInjectionTokenLimit - Math.ceil(summaryOut.length / 4)); const learnedOut = snapshot.learned && learnedBudget > 0 ? truncateByApproxTokens(snapshot.learned, learnedBudget).trim() : ""; if (!summaryOut && !learnedOut) return undefined;

return prompt.render(readPathTemplate, { memory_summary: summaryOut, learned: learnedOut, }); --- packages/coding-agent/src/memories/index.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-18 — Mnemopi retention, extraction, maintenance and retrieval; wired. Agent-end retention thresholds store previously unretained transcript segments with provenance; extraction uses user text, embeddings use separately prepared text, veracity is unknown. Fact extraction and age-gated working-to-episodic consolidation generate later-readable material. Initial prompt/pre-compaction recall selects across scoped banks, lexical/vector relevance and metadata ranking; tools request recall/full reads/edit operations. Parent and aliases share stores while automatic transcript retention is parent-owned. Source: SRC-1 packages/coding-agent/src/mnemopi/state.ts:324-369,472-551,576-606,646-700; packages/coding-agent/src/mnemopi/backend.ts:107-150; packages/mnemopi/src/core/beam/consolidate.ts:389-432,972-1059; packages/mnemopi/src/core/beam/recall.ts:731-774,963-1013.

      scope: "bank",
      extract: shouldExtract,
      extractEntities: shouldExtract,
      extractText: shouldExtract ? extractText : null,
      embedText,
      veracity: "unknown",
      memoryType: "episode",

--- packages/coding-agent/src/mnemopi/state.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

      const sleepSummary = buildSleepSummary(beam, source, chunk);
      const metadata: Metadata = { original_count: chunk.items.length, source, llm_used: false };
      if (sleepSummary.truncated) {
          metadata.truncated = true;
          metadata.original_chars = sleepSummary.originalChars;
          metadata.max_chars = sleepSummary.maxChars;
      }
      const summary = sleepSummary.summary;

--- packages/mnemopi/src/core/beam/consolidate.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

if (temporalOptions.queryEmbedding === undefined) { // Honour null (explicit "no embedding"); undefined means "derive from query text". // embedQuery() returns null when embeddings are disabled or no provider is configured, // so this is a no-op when the user has not wired one up. Float32Array → number[] // because RecallOptions exposes the narrower public shape. const derived = query.length > 0 ? await embedQuery(query) : null; temporalOptions.queryEmbedding = derived === null ? null : Array.from(derived); } let weights = normalizedRecallWeights( options.vecWeight ?? beam.config.vecWeight, options.ftsWeight ?? beam.config.ftsWeight, options.importanceWeight ?? beam.config.importanceWeight, --- packages/mnemopi/src/core/beam/recall.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-19 — Hindsight client retention and recall; wired interface. Agent-end batching/explicit retain sends transcript documents with timestamp, session metadata and tags; remote extraction/consolidation is opaque. First-turn auto-recall composes a query from incoming prompt plus recent history, calls a budgeted service, stores returned snippet and feeds prompt construction. Mental models load from operator-curated bank material with periodic refresh. Main agent also explicitly calls recall/reflect. Local clear only clears local state, not upstream memory. Alias subagents share parent bank/tool access but do not duplicate automatic recall/retention. Source: SRC-1 packages/coding-agent/src/hindsight/state.ts:293-373,425-523; packages/coding-agent/src/hindsight/backend.ts:41-147; packages/coding-agent/src/tools/memory-recall.ts:65-95; packages/coding-agent/src/tools/memory-reflect.ts:64-80.

  const history = extractMessages(this.session.sessionManager);
  const queryMessages = [...history, { role: "user" as const, content: latestPrompt }];
  const query = composeRecallQuery(latestPrompt, queryMessages, this.config.recallContextTurns);
  const truncated = truncateRecallQuery(query, latestPrompt, this.config.recallMaxQueryChars);
  const { context, ok } = await this.recallForContext(truncated);
  if (!ok) return undefined;

  this.hasRecalledForFirstTurn = true;
  if (!context) return undefined;

  this.lastRecallSnippet = context;
  return context;

--- packages/coding-agent/src/hindsight/state.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

  // Order: static instructions → mental models (stable, curated) → recall
  // (volatile per turn). Stable context first so the LLM's prior is
  // anchored on curated knowledge.
  const parts = [STATIC_INSTRUCTIONS];
  if (mentalModelsSnippet) parts.push(mentalModelsSnippet);
  if (recallSnippet) parts.push(recallSnippet);
  return parts.join("\n\n");

--- packages/coding-agent/src/hindsight/backend.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-20 — Sharpshooter friction-selected decision memory; wired implementation, prompted admission quality. Committed user-message events trigger extraction using current prompt plus at most 400 characters previous human context and 800 assistant context. Schema/evidence substring checks gate queued deltas. Scheduler or enqueue consolidates queued deltas with current files and project doctrine digest into bounded project decision files. Prompt admission requires regression, subtlety or repetition; these semantic properties are model judgments, not proven by the substring guard. Main agent later receives populated files under an injection budget; clear removes project bank. Source: SRC-1 packages/coding-agent/src/sharpshooter/backend.ts:78-129; packages/coding-agent/src/sharpshooter/extract.ts:98-148,256-293; packages/coding-agent/src/sharpshooter/consolidate.ts:56-150,202-243; SRC-2 packages/coding-agent/src/prompts/memories/sharpshooter-consolidate-system.md:5-43.

if (typeof raw.evidence !== "string" || !raw.evidence || !currentPrompt.includes(raw.evidence)) { logger.debug("Sharpshooter extraction rejected delta with unverifiable evidence", { evidence: raw.evidence }); return undefined; --- packages/coding-agent/src/sharpshooter/extract.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

              if (delta.rejectedAlternative) {
                  fields.push(`rejectedAlternative=${JSON.stringify(delta.rejectedAlternative)}`);
              }
              if (delta.rationale) fields.push(`rationale=${JSON.stringify(delta.rationale)}`);
              return `- ${fields.join("; ")}`;

--- packages/coding-agent/src/sharpshooter/consolidate.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-5 — Product editing and diagnostic feedback. Conclusion status: wired for edit/write and returned configured diagnostics; separate verification doctrine conclusion status: claimed. Trigger: model edit request; owner: edit tool and configured LSP/linter. Proposed product change is admitted through RTE-2 and edit mechanics, then reported diagnostics may inform another RTE-1 turn. Rejection/recovery includes operational errors, no-change detection and later corrective edits, not a proved universal rollback. Tool-result return and session persistence allow later read-back; selection is edited path/configured checker, with checker freshness limiting validity. Delegated workers use their granted tools; actual effect and corrective uptake remain uninspected. Guidance asks to reproduce/fix/retest, but an arbitrary diagnosis or passing check is not a supplied universal answer oracle. SRC-1 packages/coding-agent/src/edit/index.ts:607-653; SRC-2 packages/coding-agent/src/prompts/system/system-prompt.md.

return { written: request.content, diagnosticsJson: diagnostics ? JSON.stringify(diagnostics) : undefined, }; --- packages/coding-agent/src/edit/index.ts:649-652 @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-6 — Patch reviewer generation and output-shape parsing. Conclusion status: afforded for evidence-directed review through the task runtime; parser conclusion status: wired. Trigger: assigned diff; owner: reviewer model for bug/correctness judgment, renderer for structural checks. Guidance demands patch-anchored findings; parsing tests field shape, not truth of a bug. Immediate return is prioritized findings/verdict; parent or operator can accept, reject or request revision. Persistence/read-back follows task/session records, but a mandatory next-round criticism loop is uninspected. Selection is assigned diff; scope is that patch, not global product safety. No correction guarantee or answer oracle follows. SRC-1 packages/coding-agent/src/tools/review.ts; SRC-2 packages/coding-agent/src/prompts/agents/reviewer.md.

RTE-7 — Advisor note admission and delivery. Conclusion status: wired. Trigger: advisor emits note/severity; owner: emission budget/duplicate guard then session router. Proposes guidance rather than product mutation; severity and primary state choose aside, preserved message or steering. Return is delivery result; later turn sees retained note if delivered/resumed. Selection and expiry are state/budget dependent, not truth dependent. Primary model may heed or reject advice; actual uptake uninspected. Delegated visibility is the named advisor/primary channel, not all workers. No checked answer or semantic acceptance is supplied by severity. SRC-1 packages/coding-agent/src/advisor/advise-tool.ts, packages/coding-agent/src/session/session-advisors.ts.

RTE-8 — /omfg rule proposal, applicability criticism and human adoption. Conclusion status: wired. Trigger: user complaint about prior behavior. Owner: ephemeral model proposes matcher/body; code checks parsing and historical match; failed candidate plus formulated failure enters the next proposal; human can amend, cancel, choose project/global save or override a failed match. Guidance consists of complaint, prior rule and failed attempts. Test target is the rule's stated applicability to previous assistant surfaces, not corrective efficacy. Persisted OBJ-7 changes future rule availability and current registration. Immediate return is saved/rejected rule; later read-back is RTE-9; selection is human location choice and later runtime matching. Recovery is amendment/replacement/cancellation; no guaranteed undo of already-changed behavior. Delegated visibility of a newly saved rule in already-running children is uninspected. Answer oracle: assistant-history text supplies a reference for matching only; the human supplies intended behavior and adoption judgment. SRC-1 packages/coding-agent/src/modes/controllers/omfg-controller.ts:128-187,193-267, packages/coding-agent/src/modes/controllers/omfg-rule.ts; SRC-2 packages/coding-agent/src/prompts/system/omfg-user.md.

feedback: failedAttempts.length > 0 ? failedAttempts.join("\n\n") : undefined, previousRule, --- packages/coding-agent/src/modes/controllers/omfg-controller.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

await Bun.write(target.filePath, candidate.fileContent); --- packages/coding-agent/src/modes/controllers/omfg-controller.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

RTE-9 — TTSR match → interruption/reminder → continuation. Conclusion status: wired. Trigger/selector: loaded rule matches current stream/scope under configured interruption policy. Owner: coordinator; symbolic match controls stopping/injection, model determines response to guidance. Immediate effect is continuation control; saved rules/injection records support later context, with per-rule injection/cooldown policies limiting recurrence. Guarantees cover control action under matching conditions, not obedience. RTE-8 owns persistent rule revision; direct author edits are another policy source. Delegated visibility depends on each session's loaded rules and is not proven universal. No epistemic endorsement or improved capacity follows from interruption. SRC-1 packages/coding-agent/src/session/ttsr-coordinator.ts, packages/coding-agent/src/modes/controllers/omfg-controller.ts:236-267.

RTE-10 — Autoresearch candidate and benchmark execution. Conclusion status: wired for execution and recorded metrics; explanatory hypothesis formulation conclusion status: afforded. Trigger: activated experiment mode and model's run_experiment call; owner: primary model proposes a coherent change under retained playbook, fixed harness command executes it. Context: current variant, objective, baseline and notes. Product is OBJ-8; effect boundary is external project/benchmark. Immediate result: output, parsed metrics, exit/timing. Durable run feeds RTE-11 and RTE-13; delegated visibility follows explicit task/context, not a separate shared oracle. Selection is chosen variant and fixed benchmark; timeout/killed state bounds execution. A zero exit is executable completion, not proof of correctness. Answer oracle is whatever criterion/reference the supplied harness actually embodies; its validity and expected outcomes are uninspected. Mode is bounded experiment iteration, distinct from ordinary open requests. SRC-1 packages/coding-agent/src/autoresearch/tools/run-experiment.ts; SRC-2 packages/coding-agent/src/autoresearch/prompt.md.

RTE-11 — Autoresearch log, keep/discard and repository admission. Conclusion status: wired. Trigger: pending run plus agent-supplied metric/status; owner: agent judges, tool records and commits/reverts. Guidance says keep improvements while preserving correctness; implementation warns on parsed/supplied metric disagreement and scope deviation rather than vetoing solely on them. Keep on the dedicated branch commits; non-keep statuses revert, while off-branch handling is bounded by warning behavior. Run records retain lineage and discrepancy. Return is result/admission summary; later consumer uses kept code plus RTE-13 context. Stop count ends automatic mode. No independent reviewer or truth oracle is established; host/user can stop or alter the experiment setup. This admits changes to external product, not automatically oh-my-pi machinery. SRC-1 packages/coding-agent/src/autoresearch/tools/log-experiment.ts; SRC-2 packages/coding-agent/src/autoresearch/prompt.md.

RTE-12 — Autoresearch confidence arithmetic. Conclusion status: wired. Trigger: logged run, owner: computeConfidence. Input: positive unflagged current-segment metrics and best kept/baseline; output: absolute difference divided by MAD where sufficient samples/nonzero MAD exist. Stored/read by state/display and future experiment context. Selection is run/segment flags; new data replaces the statistic, not underlying history. No change admission or delegated private state is added. This is an arithmetic summary; heterogeneous variants do not independently establish repeated-condition noise or a calibrated causal probability. SRC-1 packages/coding-agent/src/autoresearch/state.ts, packages/coding-agent/src/autoresearch/tools/log-experiment.ts:200-255.

RTE-13 — Autoresearch notes/results retention and next-round prompt. Conclusion status: wired. Trigger: update_notes and before_agent_start on the active experiment branch; owner: model supplies notes, storage preserves, extension selects notes/recent runs into system prompt. Proposed changes: hypotheses/ideas/playbook text; admission requires session/tool protocol, not semantic proof. Human or model can revise notes; previous semantic truth is not guaranteed or transactionally restored. Immediate return: update result; later consumer: next experiment model invocation; persisted session notes/results span rounds and resumed experiment work. Selection uses active branch/session and recent results; scope does not establish generalization across problems. Delegated visibility beyond supplied task context is uninspected. Notes may preserve rationales or criticisms but arbitrary body content prevents inferring they did. SRC-1 packages/coding-agent/src/autoresearch/tools/update-notes.ts, packages/coding-agent/src/autoresearch/index.ts:294-405.

RTE-14 — Managed skill revision, separate from backend lesson storage. Conclusion status: wired for write/discovery; behavioral uptake conclusion status: uninspected. Trigger: manage_skill/optional learn skill payload or enabled auto-capture after sufficiently many tool calls; plan/goal mode suppress capture and automatic continuation requires opt-in. Model proposes reusable what/when/why procedure; symbolic writer sanitizes names/descriptions, checks size/body/path integrity, serializes mutation and supports create/update/delete. Authored skills take precedence in discovery. Admission is structural/filesystem and RTE-2 policy, not procedure efficacy. Lesson storage can succeed before optional skill creation fails, explicitly a partial outcome. Return is created/updated/error; future discovery returns skill metadata for model selection; current/new child visibility is configuration-dependent. Recovery: update/delete managed files; prior content rollback uninspected. Guidance recommends sparse capture and reuse; no answer oracle or mandatory content-directed criticism established. Persistence across sessions is wired availability, not demonstrated activation. SRC-1 packages/coding-agent/src/tools/learn.ts:9-138, packages/coding-agent/src/autolearn/managed-skills.ts:152-244, packages/coding-agent/src/autolearn/controller.ts:110-152, packages/coding-agent/src/extensibility/skills.ts:376-410; SRC-2 packages/coding-agent/src/prompts/system/autolearn-guidance.md.

const autoContinue = this.#settings.get("autolearn.autoContinue") === true; if (!autoContinue) return; --- packages/coding-agent/src/autolearn/controller.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

if (enabledAuthoredNames.has(capSkill.name)) continue; // an enabled authored skill owns this name --- packages/coding-agent/src/extensibility/skills.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

const bytes = Buffer.byteLength(content, "utf8"); if (bytes > MAX_MANAGED_SKILL_BYTES) { --- packages/coding-agent/src/autolearn/managed-skills.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Supporting passages for revision and checking routes

Evidence for RTE-6:

Every finding MUST be patch-anchored and evidence-backed. --- packages/coding-agent/src/prompts/agents/reviewer.md @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Evidence for RTE-5:

  • Bug fix → reproduce, fix, confirm reproduction no longer triggers. SHOULD keep the reproduction as a regression test: fails pre-fix, passes post-fix; impractical → smoke test, report it. --- packages/coding-agent/src/prompts/system/system-prompt.md @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Evidence for RTE-7:

const ADVISOR_GUIDANCE = "weigh, don't blindly obey"; --- packages/coding-agent/src/advisor/advise-tool.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Evidence for RTE-8:

  if (surface.text.length > 0 && manager.checkDelta(surface.text, surface.context).length > 0) {
      matches.push(surface);
  }

--- packages/coding-agent/src/modes/controllers/omfg-rule.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Evidence for RTE-8:

                  "Couldn't confirm this rule matches the conversation. Save anyway?",

--- packages/coding-agent/src/modes/controllers/omfg-controller.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Evidence for RTE-10:

      const passed = execution.exitCode === 0 && !execution.killed;

--- packages/coding-agent/src/autoresearch/tools/run-experiment.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Evidence for RTE-11:

  1. Keep the primary metric as the decision maker:
  2. keep when it improves;
  3. discard when it regresses or stays flat;
  4. crash when the run fails;
  5. checks_failed when validation fails (you decide what validation means; run it through the regular bash tool). --- packages/coding-agent/src/autoresearch/prompt.md @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a
      const metric = params.metric;
    

    --- packages/coding-agent/src/autoresearch/tools/log-experiment.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

              `Logged metric ${metric} differs from parsed primary ${pendingRun.parsedPrimary}. Both values stored.`,
    

    --- packages/coding-agent/src/autoresearch/tools/log-experiment.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Evidence for RTE-12:

return Math.abs(bestKept - baseline) / mad; --- packages/coding-agent/src/autoresearch/state.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Evidence for RTE-13:

              has_notes: state.notes.trim().length > 0,
              notes: state.notes,

--- packages/coding-agent/src/autoresearch/index.ts @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

Memory route return, revision and recovery audit

This audit supplements the source-native route records, using their cited sources and the specialist's write/read-back findings. All local operation conclusions are wired unless an alternative or unknown is stated; actual model activation is uninspected throughout.

Route Immediate return and later consumer Delegated visibility and selection Admission, expiry, rejection and recovery Guidance and retained content
Audit of RTE-15 Append/load/branch result; later model gets selected history Main or explicitly resumed child; leaf/parent IDs select actual entries Branch selection changes reliance without deleting all history; malformed/corrupt storage recovery beyond inspected code uninspected; no semantic veto Conversation and selected branch context; unchanged retention is not a separate theory revision
Audit of RTE-16 Maintenance result; next context rebuild replays checkpoint/tail Per-session consumer; checkpoint/tail IDs and budget; child-specific application depends on child session Method fallback, payload budgets, kept tail, shake recovery references; no semantic preservation guarantee; remote recovery internals uninspected Continuation goals/decisions and requested rationale; prose, mechanical archive and opaque alternatives remain distinct
Audit of RTE-17 Extraction/consolidation/write result; later project model/advisor gets summary and lessons; agent requests files Current project, idle prior files excluding active thread; summary budget before lesson budget; detailed files pulled by agent Empty/schema-invalid/model-error extraction rejected; redaction and job ownership; age/idle bounds; regeneration replaces generated files and can lose manual edits; learned.md separate Trace-derived heuristics/lessons, optional why/context and generated procedure files; memory advice explicitly yields to current repository/user evidence
Audit of RTE-18 Retain/edit/recall result; later bank-scoped model Aliases share stores/tool access; parent owns auto-retention; lexical/vector/rank metadata selects results Unknown transcript veracity, eligibility rules for update/forget/invalidate; age degradation and vector invalidation; advanced maintenance limits explicit Extracted facts and transcript summaries; source IDs/summary_of preserve lineage, not necessarily explanatory rationale
Audit of RTE-19 Queued retain/service recall/reflect response; later main model Parent owns auto-retain/recall; aliases share bank tools; first-turn query/history/tags/budget inputs, opaque service selector Client errors limit delivery; periodic mental-model refresh; local clear does not withdraw upstream memory; remote rejection/expiry uninspected Imported service text and curated bank material; upstream transformation and warrant uninspected
Audit of RTE-20 Delta enqueue/consolidation result; later project model receives decisions Project files and injection budget select populated instruction blocks; child-specific fresh visibility uninspected Evidence must occur in current prompt; model judges friction/importance; consolidation replaces bounded files, clear withdraws project bank User-message evidence, optional rationale/rejected alternatives read by consolidator; final reason retention selective, no independent semantic validation

For CMP-5, CMP-6 and CMP-7, inference use conclusion status: wired. Deployed immutable version identity conclusion status: uninspected; provider-side parameter changes conclusion status: uninspected. The inspected calls produce text/vectors, not an observed training intervention. For CMP-8, client inference/service use conclusion status: wired; actual upstream component identity and parameter changes conclusion status: uninspected. This preserves fixity as a distinct question from knowledge retention.

Claims

CLM-1 — Source-native positioning: coding agent with IDE capabilities. Conclusion status: claimed. RTE-1, RTE-2 and RTE-3 establish runtime wiring; they do not establish superiority to alternatives. SRC-2 README.md:1-32.

A coding agent with the IDE wired in. --- README.md @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

CLM-2 — README advertises edit-format performance gains. Conclusion status: claimed. Its table is source-reported marketing/design evidence, not an inspected run or causal experiment in this source boundary. The linked external blog and benchmark artifacts were not included; numerical or component-effect attribution is uninspected. SRC-2 README.md:110-129.

Tenfold lift the moment the edit format stops eating the model alive. --- README.md @ be6cb8217cd4c1dafcc86793ae5d809ea4d7396a

CLM-3 — Reviewer doctrine requires evidence-backed patch findings and a correctness verdict. Conclusion status: claimed; RTE-6 establishes only the bounded generation/parser route, not an independent shipping oracle. SRC-2 packages/coding-agent/src/prompts/agents/reviewer.md.

CLM-4 — TTSR is intended to correct behavior by matching, interruption and retry. Conclusion status: claimed; RTE-8 and RTE-9 establish control and adoption wiring, not successful correction. SRC-2 README.md; SRC-1 coordinator on RTE-9.

CLM-5 — Autoresearch doctrine asks for autonomous metric-guided iteration preserving correctness. Conclusion status: claimed; RTE-10, RTE-11, RTE-12 and RTE-13 establish execution, agent-decided disposition and recurrence, not compliance with every methodological requirement. SRC-2 packages/coding-agent/src/autoresearch/prompt.md.

Evidenced absences

ABS-1 —No observed memory-faithfulness result in the commissioned evidence; absent. Input explicitly supplies source-only evidence with no observed runs or causal experiments. No execution probe was run. This supports faithfulness_tested no only for this analysis evidence, not a claim that the project has never tested recall.

No additional runtime absence is asserted. Uninspected provider/deployment properties remain limitations.

Behavioral-authority paths

BAP-1 — Tool policy → wrapper → executor. Consumer: registered-tool dispatch; channel: symbolic resolved policy; force: enforcement of allow/deny/prompt; horizon: one call. Operational permission, no epistemic warrant. Conclusion status: wired; see RTE-2.

BAP-2 — Model call/tool result → next agent turn. Consumer: chat model; channel: conversation context; force: advisory knowledge and user/task instruction, with symbolic dispatch controlling effects; horizon: current loop and separately retained session. Delivery is wired; activation uninspected. See RTE-1.

BAP-3 — Child status/patch applicability → parent worktree apply-back. Consumer: isolation runner; channel: result fields and Git checks; force: operational admission; horizon: one delegated change set. No correctness certification. Conclusion status: wired; see RTE-3.

BAP-4 — Extension module → live runtime. Consumer: factory loader/tool and hook registry; channel: executable registration and direct API; force: execution and routing; horizon: loaded session and future reloads from retained configuration. Conclusion status: wired; see RTE-4.

BAP-5 — Session/checkpoint → main model/provider context; automatic identifier-selected delivery, advisory conversation force across resumed turns. Conclusion status: wired; see RTE-15, RTE-16.

BAP-6 — Project summary/lessons → main model and supported advisors; bounded prompt block and requested memory-file reads, advisory knowledge force across project sessions. Conclusion status: wired; see RTE-17.

BAP-7 — Mnemopi rows/metadata → recall selector and model; bank routing and ranking select advisory content, while explicit tools pull rows. Conclusion status: wired; see RTE-18.

BAP-8 — Hindsight bank → client prompt/tool response → model; service-selected advisory material and curated mental models, across bank/session horizon. Client delivery conclusion status: wired; upstream warrant uninspected. See RTE-19.

BAP-9 — Sharpshooter decisions → main system prompt; project-scoped instruction unless user overrides, across sessions until file revision/clear. Conclusion status: wired; compliance uninspected. See RTE-20.

BAP-10 — Review/advisor → parent model/operator; report, aside or steering channels; advisory criticism whose severity changes timing, not truth. Current/resumed turn horizon. Advisor routing conclusion status: wired. Review reliance conclusion status: afforded; see RTE-6, RTE-7.

BAP-11 — Saved TTSR rule → matcher/coordinator → main model; symbolic interruption plus instruction injection; session and loaded-rule horizon. Conclusion status: wired; see RTE-8, RTE-9.

BAP-12 — Agent experiment disposition → Git commit/revert; operational enforcement on dedicated branch; next-variant horizon. Conclusion status: wired; see RTE-11.

BAP-13 — Experiment notes/results → next model system prompt; instruction and evidence, not endorsement; ongoing/resumed experiment horizon. Conclusion status: wired; see RTE-13.

BAP-14 — Managed skill metadata/body → future skill discovery and selected model context; procedural instruction, lifetime until update/delete/shadowing; discovery conclusion status: wired, actual uptake conclusion status: uninspected. See RTE-14.

Runtime account

The shipped omp binary enters the coding-agent CLI; main selects interactive, print, RPC/RPC-UI or ACP paths. The SDK resolves the model, settings, context and tool registry, builds the Agent and AgentSession, and restores existing messages where present. SRC-2 packages/coding-agent/package.json:28-32; SRC-1 packages/coding-agent/src/main.ts:1866-1880,2045-2118, packages/coding-agent/src/sdk.ts:3531-3645. The operator or runtime client supplies the problem; provider credentials are resolved at dispatch. The source commit fixes code and shipped prompts but does not fix remote model weights or a user's active configuration.

RTE-1 owns progression. The model selects content/tool calls within supplied tools, while code handles argument preparation, context transformations, scheduling, control queues, stop reasons and aborts. Results return to subsequent turns, so an execution error can inform another attempt without thereby becoming a formulated criticism of a theory. The terminal product is streamed text/structured output plus external effects and retained session material. RTE-2 determines permission at registered dispatch; it is not a correctness gate. RTE-3 supplies child control and optional change integration; RTE-4 admits executable capabilities.

Material alternate paths are separately bounded: direct SDK/lower-level agent clients can supply tools and stream functions; RPC and ACP introduce host clients; provider-native execution bridges and in-band tool formats alter dispatch representation; eval/Python kernels, shell access and MCP tools introduce other executors; extensions can execute directly; task workers create additional sessions and may use isolated working directories. The SDK registry wrapper covers registered tools, including bridge tools that use that registry. That evidence cannot be extended to arbitrary extension code, child processes' internal actions, provider-hosted tools or excluded host implementations. SRC-1 packages/coding-agent/src/sdk.ts:3531-3617, packages/agent/src/agent-loop.ts:1593-1640, packages/coding-agent/src/extensibility/extensions/loader.ts:277-290. These alternatives limit guarantees; they are not assumed to be equivalent deployments.

Four static forcing cases challenge the ordinary route:

Case Boundary and finding Evidence and conclusion prevented
Hook rewrites an allowed tool call RTE-2 resolves policy against effective arguments after interception; required approval without UI rejects SRC-1 wrapper excerpts on RTE-2; wired admission protection, not observed security isolation
Delegation from an interactive parent RTE-3 snapshots settings but switches child mode to yolo; explicit deny policies still apply SRC-1 executor and approval resolver; parent approval authorizes headless work, not a repeated per-action human veto
Isolated child's patch conflicts or its run fails RTE-3 distinguishes successful run from applicability, preserves failed application artifacts and reports manual resolution SRC-1 isolation-runner excerpts; Git applicability is not semantic acceptance
Retained context exceeds a budget or changes branch Memory records trace compaction/checkpoints and branch/context selection rather than treating all stored bytes as delivered Specialist-integrated records below; delivery and preservation depend on the selected route, with opaque alternatives explicit

Execution disposition: no dynamic check planned. A mocked approval rewrite, isolated patch-apply fixture and live memory/compaction exercise were considered. Static branches establish the conditional wiring needed here. Mocked executions would not establish deployment containment or learning; a live run would additionally require a configured runtime, provider credentials and external services and would not by itself establish benefit. No target execution, package installation or credential inspection was attempted. Accordingly no outcome is labelled observed or causally supported.

The capability surface is broad, the current grant set is uninspected, and the deployed isolation envelope is uninspected. The source distinguishes tool policy, allowed tools, worker working directory and provider/host contracts; this analysis preserves those distinctions. Recovery mechanisms preserve or return state selectively, without a global transaction over file writes, commands and remote effects.

For ordinary open requests, the model proposes edits and decides what to attempt next; symbolic tools expose results; the operator can supply instructions or veto gated calls. Task and Git mechanisms decide operational application. A task may bring tests/reference outcomes, but no universal answer oracle is supplied by the generic runtime. The dedicated rule-revision and experiment routes below have their own proposers, evaluators, admission and retention mechanisms. Their computational and human contributions must not be generalized to all requests.

Lens scoping

Memory/context scope

Full depth. Trigger evidence: SRC-1 session persistence, compaction, backend interfaces and memory tools; SRC-2 memory documentation and prompts. The scope is accumulated or changed retained objects and their actual later consumers, including material backend/compaction alternatives. Current-run OBJ-2 and static shipped guidance are context but do not become read-back simply by appearing in a prompt. Exact adopted scope and records are those of the integrated specialist profile. Full depth is warranted because context reconstruction and retained memory are material runtime responsibilities.

Epistemic scope

Full depth. Trigger evidence: SRC-1 product mutation/diagnostics, review, complaint-to-rule generation, stream-rule application and experiment admission; SRC-2 review/revision doctrine and CLM-2 performance assertions. RTE-1, RTE-2, RTE-3 and RTE-4 establish the runtime and operational authority against which the epistemic overlay is interpreted. Memory transformations are integrated from the specialist rather than independently reclassified. Provider reasoning, external service internals and actual candidate histories are uninspected; no whole-system completeness of warrant or successful learning is inferred.

Lens outputs

Memory/context lens

The fresh specialist inventoried OBJ-12, OBJ-13, OBJ-14, OBJ-15, OBJ-16 and OBJ-17 and RTE-15, RTE-16, RTE-17, RTE-18, RTE-19 and RTE-20. Adopted conclusions are implementation wiring or afforded alternatives, not observed activation. The frontmatter profile is intentionally bounded to those session/backend surfaces; managed-skill, TTSR, /omfg and autoresearch routes remain separately described outside its aggregate.

Write side

Local rollout processing is staged, bounded and best-effort. Defaults scan up to 300 threads, process up to 64 rollouts per startup, cap age at 30 days and require 12 hours idle; phase-one input defaults to 4,000 approximate tokens, consolidated raw/summary input is capped separately at 20,000/12,000. Invalid JSON/schema, model errors and empty extracted memory reject output. Secrets are redacted. Stage two rewrites MEMORY.md and summary, replaces generated skills and prunes omitted files. It does not read hand-edited prior MEMORY.md into consolidation: its model input is raw memories plus rollout summaries. Thus a hand edit is not guaranteed to survive regeneration. SQLite ownership/watermark jobs coordinate work; they are execution metadata, not evidence of semantic truth. RTE-17; SRC-1 packages/coding-agent/src/memories/index.ts:60-77,724-807,872-1009.

Learn captures a durable lesson and may separately create/update a managed skill; partial skill failure can occur after memory was already stored. Local learned.md remains separate from regenerated files and retains up to 100 lessons. Its optional context can carry a reason, and read-back injects the resulting line; neither freeform field guarantees a useful explanation. The learn schema asks for what/when/why. Parent integration must include managed-skill execution/discovery if it broadens this memory profile; this report does not infer adoption from a generated file. SRC-1 packages/coding-agent/src/tools/learn.ts:9-27,51-138; packages/coding-agent/src/memories/index.ts:1284-1297,1338-1364.

Mnemopi separates transcript retention from user-only fact extraction, records unknown veracity for transcripts and tool veracity for learn. Memory editing is store-sensitive: fact projections are not editable, working rows can be updated/forgotten, eligible rows can be invalidated through Beam. Age-based degradation truncates old episodic content and invalidates stale vectors; sleep preserves source-row lineage and can later be recalled. This is durable curation, while recall-result deduplication alone is selection rather than retained-memory deduplication. Optional advanced operations remain uninspected. RTE-18; SRC-1 packages/coding-agent/src/mnemopi/state.ts:324-369,522-551,686-700; packages/mnemopi/src/core/beam/consolidate.ts:389-432,835-896,972-1059.

Sharpshooter retains optional rationale and rejected alternatives in the delta queue; its consolidator explicitly reads those fields. Final instructions preserve reasons only selectively: doctrine asks for a rejected alternative when it remains a live temptation. Therefore the consolidator has a wired reason-consuming route, and the main agent may receive reasons retained in a bullet, but the final files do not guarantee a complete reason history. Consolidation does not independently test whether a decision is correct. RTE-20; SRC-1 packages/coding-agent/src/sharpshooter/consolidate.ts:56-78; SRC-2 packages/coding-agent/src/prompts/memories/sharpshooter-consolidate-system.md:34-43.

Continuation summaries explicitly request decision rationale, later read by the resumed model. That is a reason-retaining design, not evidence a generated summary actually included the right reason. Raster archives can preserve original rationale if it survives layout/budget pruning. Shake offloads original reasons with other heavy content, leaving a recovery reference. Opaque remote compaction prevents assessing reason retention. Mnemopi source IDs and summary_of supply lineage rather than an explanation of a prescription; no mandatory rationale field was found in the inspected extraction/storage path. Hindsight upstream rationale handling is unknown. RTE-16, RTE-18, RTE-19.

Read-back

RTE-15 and RTE-16 push context based on selected leaf and retained checkpoint IDs. Kept-tail boundaries select actual delivered entries, not merely catalog metadata. The user-facing transcript and provider replay deliberately differ for remote replacement history. A visible summary alone cannot tell what the provider consumes. Snapcompact reattaches image blocks on context rebuild, with oldest images omitted when payload budget is exceeded; text edges and omission notices remain. Shake references afford an agent recovery read. No model reading success was measured.

RTE-17 pushes the current project's summary first and lessons only from remaining injection budget (default shared limit 5,000 approximate tokens). Summary can consume the whole allowance, leaving learned lessons retained but undelivered. Session-file keyed cache also means a write is not necessarily visible immediately. Main agent is explicitly instructed to read MEMORY.md/skills if needed via memory://root; this is a named pull consumer, not merely filesystem availability. Current files and user instructions supersede this advice. SDK prompt rebuilding also shares the memory block with advisors; tool access caveats remain. SRC-1 packages/coding-agent/src/memories/index.ts:182-228,277-293; packages/coding-agent/src/sdk.ts:3098-3106.

RTE-18 and RTE-19 add automatic first-turn/pre-compaction recall and agent-requested recall/reflect. Mnemopi selects project/global bank visibility, merges scoped results, clips by recall limit and renders under an injection budget. Local lexical, embedding, importance, time and veracity scores affect rank; Hindsight server selection is opaque even though query/tags/budget are explicit. Mnemopi memory://id full reads let the agent inspect a clipped result before editing. Hindsight explicitly rejects that addressing and directs the agent back to recall/reflect. Aliases share the parent's state but do not independently auto-retain internal exploration transcripts. SRC-1 packages/coding-agent/src/mnemopi/state.ts:411-494; packages/coding-agent/src/mnemopi/backend.ts:107-150; packages/coding-agent/src/hindsight/backend.ts:41-101; packages/coding-agent/src/internal-urls/memory-protocol.ts:302-335.

RTE-20 pushes populated current-project decision files at prompt rebuild. This is coarse project-scoped delivery capped by injectionTokenLimit, not query-specific retrieval; the lexical search method is a separate requested route and cannot establish lexical push. Final instructions carry stronger authority than local advisory memory, but their admission quality still depends on generated judgment.

Comparison rationale

Known unions cover the scoped interface and object boundary, not excluded service internals. Storage includes explicit SDK alternatives at the weaker afforded basis; remote service-object does not guess its database. Rank metadata and branch/bank routing are operative memory parts. Manual editing and automatic transformation coexist. Static prompts are evidence of consumer authority, not counted as accumulated memory themselves.

Trace learning is yes because multiple trace-fed durable summaries/facts/guidance objects are wired to later consumers. It neither means parameter training nor proves improvement. Sources include stored sessions/tool outcomes and committed message events. A summary that helps continue work still qualifies. Trace horizons/timing/form remain uncertain across all alternatives rather than being inferred from session IDs or visible prose. Raster language is human-language content transmitted through images, but the three-value representation vocabulary lacks a clean raster token; encrypted checkpoint payload is additionally opaque.

Unknown curation and lineage axes are deliberately conservative: imported/custom history, optional advanced Mnemopi behavior and upstream transformations have not all been characterized. The known local operations are recorded in the write-side account and should survive integration even where the aggregate assessment stays uncertain. External service internals are excluded, but their visible interface remains included. No inference from the name reflect establishes LLM synthesis for Mnemopi.

Epistemic lens

1. Source-and-claim boundary

See SRC-1 and SRC-2, CLM-3, CLM-4 and CLM-5. Assessed routes: product diagnostics, review/advice, rule revision and experiment execution/admission. Excluded checker/provider internals and missing run evidence prevent behavioral efficacy or causal attribution. Memory transformation warrant is annotated separately below. Unassessed debugger/browser/security/plugin evaluators prevent a system-complete negative.

2. Epistemic-object annotations

See OBJ-1 and OBJ-4, OBJ-5, OBJ-6, OBJ-7, OBJ-8, OBJ-9, OBJ-10 and OBJ-11 for identity, form and sources. Product adequacy/reviewer claims and proposed optimization efficacy are ampliative candidates; diagnostics are imported checker reports. Rules/skills prescribe behavior without a required truth-apt efficacy claim. Notes/advice remain indeterminate without an instance. Core tool calls and registry settings have no candidate truth-apt output merely by being executable; no lifecycle record is assigned to OBJ-2 or OBJ-3.

3. Authority-route ledger

Each row annotates an existing canonical route with one epistemic function. Architectural status defaults to implemented; explicitly limited doctrine rows retain that weaker status. Possible results are not observed outcomes.

Proposal / function / object / content relation Target, evaluator, activation and result Authority, consumer/channel/horizon; evidence and limit
RTE-5; check/evidence production; OBJ-1 → OBJ-4; acquisition/import Edited file; configured LSP/linter after write when enableLsp and diagnosticsOnEdit; report or no diagnostics Advisory evidence through edit tool result to worker for later correction; warrants only the configured checker report at its freshness boundary. No successful-build/product acceptance follows. packages/coding-agent/src/edit/index.ts, createEditWritethrough; packages/coding-agent/src/lsp/diagnostics.ts, getDiagnosticsForFile
RTE-5; check/evidence production; OBJ-1; no content change; doctrine only in this overlay Bug fix: reproduce, fix, confirm; feature: exercise new behavior; worker instructed before final delivery Instruction-level completion criterion, intended scope changed behavior. Parent may supply implemented shell/core loop. packages/coding-agent/src/prompts/system/system-prompt.md, Verify (SRC-2). No inspected run proves compliance
RTE-6; content transformation; OBJ-5; ampliative conjecture Assigned patch and consumer context; reviewer model instructed to report only provable impact, patch introduction and no unstated assumptions Candidate bug/correctness claims to parent/operator via structured yield; epistemic license is model assessment under doctrine, not verified program truth. packages/coding-agent/src/prompts/agents/reviewer.md (SRC-2); doctrine only for generation in this overlay, task runtime owned by parent. CLM-3
RTE-6; check/evidence production; OBJ-5; no content change Finding payload; parseFindingDetails checks strings, priority, finite confidence within 0–1, file path and numeric lines before render parsing Accepts renderable field structure or undefined; does not check cited file, patch overlap or impact. Consumer reviewer render path, display horizon current result. packages/coding-agent/src/tools/review.ts, parseFindingDetails. CLM-3
RTE-7; operational admission/selection/consumption; OBJ-6; no content change Emitted note/severity and primary state; routing selects aside/preserve/steer Aside waits; concern/blocker normally interrupt; final handoff exception lets blocker wake primary; paused states preserve. Model is told to weigh advice. Epistemic acceptance absent: emission guard tests budget/duplication, not truth. Consumer primary via advisory tagged custom message, current/resumed turn. packages/coding-agent/src/advisor/advise-tool.ts, resolveAdvisorDeliveryChannel/ADVISOR_GUIDANCE; packages/coding-agent/src/session/session-advisors.ts, acceptAdvice/routeAdvice
RTE-8; content transformation; OBJ-7; non-truth-apt policy/content update: proposed matcher and corrective instructions User complaint plus history and failed attempts; ephemeral model generates candidate, up to three attempts, user amendments restart Proposes policy; no warrant that correction will work. Consumer rule parser then operator; draft horizon. packages/coding-agent/src/modes/controllers/omfg-controller.ts, generateCandidate; packages/coding-agent/src/prompts/system/omfg-user.md (SRC-2)
RTE-8; check/evidence production; OBJ-7; no content change Rule's ability to reach an assistant-history surface; TtsrManager matcher plus scope feedback; each candidate and repaired escaping tested Match/no-match plus textual failure feeds next proposal. License: trigger matches history within scope; not identification of the complained-of error, future precision, correction efficacy, or transfer. Consumer generator/operator; current generation cycle. packages/coding-agent/src/modes/controllers/omfg-rule.ts, validateRuleAgainstAssistantHistory; packages/coding-agent/src/modes/controllers/omfg-controller.ts, generateCandidate
RTE-8; disposition/acceptance; OBJ-7; no content change Proposed policy; human reviews visible candidate, chooses project/global save, amendment or cancellation; failed match can be overridden Operational adoption of instruction, not truth-apt acceptance. Save choice permits persistence and live registration. Consumer saveCandidate; user-selected project/global horizon. packages/coding-agent/src/modes/controllers/omfg-controller.ts, runRequest/saveCandidate
RTE-8; retention; OBJ-7; no content change Human-selected target; Bun.write persists rule; overwrite asks human Retains user-adopted policy; no epistemic license added. Consumer future rule discovery, filesystem channel; project/global lifetime. packages/coding-agent/src/modes/controllers/omfg-controller.ts, saveCandidate/resolveTarget
RTE-9; behavior/policy adaptation; OBJ-7; non-truth-apt policy/content update: interrupt/inject/retry Saved rule registered live; matching stream under configured context/scope Enforcing interruption where applicable, then guidance injection and agent.continue; content compliance remains model-dependent. Consumer primary, custom system reminder, session and future loaded-rule horizon. packages/coding-agent/src/modes/controllers/omfg-controller.ts, registerLive; packages/coding-agent/src/session/ttsr-coordinator.ts, handleMatches. CLM-4
RTE-10; content transformation; OBJ-8; ampliative conjecture; doctrine only for proposed causal/adequacy content Primary instructed to identify bottleneck, establish baseline, make coherent experiment A code variant is candidate solution; generated hypothesis need not exist beyond description. Consumer benchmark and next iteration. packages/coding-agent/src/autoresearch/prompt.md (SRC-2); parent owns general edits. CLM-5
RTE-10; check/evidence production; OBJ-8 → OBJ-9; acquisition/import Fixed DEFAULT_HARNESS_COMMAND, external benchmark environment; run_experiment captures output, parses METRIC/ASI, records exit/timing passed = zero exit and not killed, not semantic correctness. Parsed metric licenses only what harness actually measures; no harness-validity check. Consumer agent via result and log, run horizon. packages/coding-agent/src/autoresearch/tools/run-experiment.ts, execute. CLM-5
RTE-11; disposition/acceptance; OBJ-8 / OBJ-9; no content change Agent calls log_experiment with status and metric after a pending run; doctrine criterion improve primary metric, preserve correctness Operational keep/discard/crash/checks_failed under agent judgment. Implementation does not enforce improvement, correctness, agreement with parsed metric, or justified scope. Epistemic acceptance criterion is doctrine, not mechanically established; intended use next iteration baseline. packages/coding-agent/src/autoresearch/tools/log-experiment.ts, execute; packages/coding-agent/src/autoresearch/prompt.md (SRC-2). CLM-5; mismatch: metric/justification checks warn rather than veto
RTE-11; retention; OBJ-9; no content change log_experiment markRunLogged records supplied status/metric, commit, deviations; parsed values already retained Preserves lineage including discrepancies, not independent endorsement. Consumer experiment state and operator, durable session records. packages/coding-agent/src/autoresearch/tools/log-experiment.ts, markRunLogged; packages/coding-agent/src/autoresearch/tools/run-experiment.ts, markRunCompleted
RTE-11; operational admission/selection/consumption; OBJ-8; no content change Supplied keep on dedicated branch commits modified files; other statuses revert; off-branch keep warns and leaves edits Enforcing repository transition based on agent disposition; later variants start from kept code. Not automatically lifecycle integration of an accepted causal claim. packages/coding-agent/src/autoresearch/tools/log-experiment.ts, keep branch/revertFailedExperiment. CLM-5
RTE-12; content transformation; OBJ-9; entailed derivation Unflagged positive metrics in current segment; computeConfidence computes abs(best kept − baseline)/MAD, at least three values and nonzero MAD Arithmetic summary under code semantics; no calibrated probability or causal identification. Values mix candidate variants, so MAD is not independently established measurement noise. Consumer state/display/agent, current segment. packages/coding-agent/src/autoresearch/state.ts, computeConfidence
RTE-13; retention; OBJ-10; no content change to supplied body, non-ampliative append formatting Active session; update_notes accepts body or appends idea and updates session storage No semantic acceptance; content may be untested hypothesis or policy. Consumer session state, durable playbook. packages/coding-agent/src/autoresearch/tools/update-notes.ts, execute/appendIdea
RTE-13; operational admission/selection/consumption; OBJ-9 / OBJ-10; no content change Active autoresearch branch at before_agent_start; notes, recent runs, flags, unjustified deviations supplied into system prompt Can inform criticism/revision and future selection, but actual uptake unobserved. System-prompt placement gives retained prose behavioral authority without epistemic endorsement. agent_end can schedule continuation; max iteration cap turns mode off. packages/coding-agent/src/autoresearch/index.ts, before_agent_start/agent_end; packages/coding-agent/src/autoresearch/tools/log-experiment.ts, maxExperiments. CLM-5

4. Per-object lifecycle disposition

No candidate instance or candidate-linked trace was observed for any object. Every lifecycle phase below has observed candidate state no instance observed, including phases whose implementation exists. This is not a failed or suspended run.

  • OBJ-1, ampliative candidate adequacy: observation from source/request is parent-owned; conjecture parent-owned; consequence derivation not determinable; check RTE-5 implemented and RTE-5 doctrine only in this overlay; acceptance and post-acceptance lifecycle integration not determinable. No checker-clean outcome establishes full requested behavior. Missing: candidate, expected behavior, executed reproduction/validation and evidence-consuming decision.
  • OBJ-5, ampliative bug/correctness claims: observation/conjecture through RTE-6 doctrine only in this overlay; consequence tracing demanded by reviewer doctrine; evidence requirements doctrine only; RTE-6 implemented structural check. Acceptance for merge and post-acceptance integration: no route found in scoped reviewer prompt + shape/parser + launcher reads; parent/operator may act elsewhere. Intended reliance is patch assessment, not a source-proven shipping oracle. Missing candidate-linked support and consequential acceptance.
  • OBJ-8, ampliative optimization candidate: observation and conjecture RTE-10 doctrine only; derived consequence (expected metric improvement with preserved correctness) doctrine only, no enforced explicit prediction; test RTE-10 implemented; disposition RTE-11 implemented as agent judgment under improvement/correctness doctrine, criterion compliance unverified; post-acceptance integration not determinable. RTE-11 is implemented operational use and RTE-13 is implemented evidence reuse; neither alone establishes epistemic acceptance. Missing actual run, valid benchmark domain, controlled contrast, acceptance rationale and later uptake.
  • OBJ-4, non-ampliative acquisition/import: RTE-5; discovery lifecycle not applicable to imported checker report. Warrant remains checker/version/file-domain limited; checking internals and applicability proof unassessed.
  • OBJ-9, non-ampliative acquisition plus arithmetic derivation: RTE-10, RTE-11, RTE-12, RTE-13. Discovery lifecycle not applicable to captured values/arithmetic. Supplied status/metric is agent judgment, not entailed by captured output. Provenance distinguishes them; no correctness or causal inference passes through merely by storage.
  • OBJ-6, indeterminate: RTE-7. A note can repeat evidence, conjecture a bug or prescribe action; no instance decides. Criticism can change attention/continuation before independent acceptance. Need actual note, source observations and primary response to classify content and uptake.
  • OBJ-10, indeterminate: RTE-13, RTE-13. Could summarize, hypothesize or prescribe; retained lineage is session context but no required per-claim evidence link. Need actual old/new note and run provenance to decide preservation versus ampliation. Prompt reuse is not lifecycle integration absent acceptance.
  • No lifecycle record for OBJ-7: no required candidate truth-apt output for this object; direct policy routes RTE-8, RTE-8, RTE-8, RTE-8, RTE-9. Matcher applicability is tested, but the prescriptive body is not thereby an accepted efficacy claim.

No lifecycle record for OBJ-11: managed procedure text is a policy proposal here, with RTE-14 operational admission and availability, not demonstrated truth-apt efficacy acceptance.

Memory objects OBJ-12 and OBJ-13 preserve/acquire or reshape traces; provider payload content is indeterminate. OBJ-14 and OBJ-15 contain model-extracted claims/guidance whose preservation versus ampliation cannot be settled from schemas alone. OBJ-16 has an opaque service producer; OBJ-17 is prescriptive decision text with imported user evidence. RTE-15 and RTE-16 retain/deliver context; RTE-17 and RTE-18 retain transformed candidates; RTE-19 imports service results; RTE-20 gates evidence substrings and prompted friction judgments before instruction storage. Architectural status: implemented for the inspected client/local paths; observed candidate state: no instance observed for all these objects. No record of epistemic acceptance or post-acceptance lifecycle integration is inferred from storage/prompting. A source substring licenses provenance to the prompt, not truth of the proposed decision; unknown veracity, advisory verification instructions and opaque service provenance retain their limits.

5. System-claim versus route comparison

Claim Doctrine/source Implementation Observed / causal support Supported conclusion; mismatch/unknown
CLM-3 SRC-2 README.md section 10, verdict whether change ships; reviewer prompt specifies correctness excluding style/nits RTE-6 (generation doctrine in this overlay), RTE-6; parent supplies task dispatch None Explicit model verdict and prioritization are intended outputs; correctness or shipping safety not independently established by shape checking
CLM-4 SRC-2 README.md section 04 says rule-triggered interruption/retry yields course correction RTE-8, RTE-8, RTE-9 None; README capture descriptions are doctrine/attributed illustration, not inspected run artifacts Mechanism changes prompting/continuation. Cannot establish that correction works or persists behaviorally. Compaction survival detail belongs to memory worker
CLM-5 SRC-2 packages/coding-agent/src/autoresearch/prompt.md says autonomous loop, improve primary metric, preserve correctness, confidence relative to observed noise RTE-10, RTE-10, RTE-11, RTE-11, RTE-11, RTE-12, RTE-13, RTE-13 None Experiment execution and mechanical commit/revert/reuse implemented. Decision remains agent-supplied. Confidence formula summarizes variation across runs; source does not establish repeated constant-condition noise measurement or causal component isolation

CLM-1 is operational positioning supported by RTE-1/RTE-3 wiring, with superiority unassessed. CLM-2 remains claimed performance: no inspected run or causal support. No universal knowledge-production claim was established within the assessed doctrine.

6. Bounded conclusion

Product diagnostics supply domain-limited feedback; review and advice formulate candidates that can guide later work. Rule applicability checks govern a narrow content property and human adoption. Experiments enforce repository consequences of model decisions, with separate metrics and preserved discrepancies. None of these permissions or storage transitions certifies a model's substantive claim. Theory-builder and learning findings are stated separately in Bounded synthesis.

Reconciliation

Memory input/source/method identities and complete report status were verified. The specialist report is provenance; all adopted mechanisms, quotes, scope and uncertainty are retained here. Proposal mapping:

  • MEM-OBJ-1 → OBJ-12.
  • MEM-OBJ-2 → OBJ-13.
  • MEM-OBJ-3 → OBJ-14.
  • MEM-OBJ-4 → OBJ-15.
  • MEM-OBJ-5 → OBJ-16.
  • MEM-OBJ-6 → OBJ-17.
  • MEM-RTE-1 → RTE-15.
  • MEM-RTE-2 → RTE-16.
  • MEM-RTE-3 → RTE-17.
  • MEM-RTE-4 → RTE-18.
  • MEM-RTE-5 → RTE-19.
  • MEM-RTE-6 → RTE-20.
  • MEM-CMP-1 → CMP-5.
  • MEM-CMP-2 → CMP-6.
  • MEM-CMP-3 → CMP-7.
  • MEM-CMP-4 → CMP-8.
  • MEM-ABS-1 → ABS-1.

Epistemic proposal mappings (several functional rows annotate one source-native pipeline; canonical IDs are allocated once and never reassigned):

  • EPI-OBJ-PRODUCT → OBJ-1.
  • EPI-OBJ-DIAGNOSTICS → OBJ-4.
  • EPI-OBJ-REVIEW → OBJ-5.
  • EPI-OBJ-ADVICE → OBJ-6.
  • EPI-OBJ-RULE → OBJ-7.
  • EPI-OBJ-EXPERIMENT → OBJ-8.
  • EPI-OBJ-RUN → OBJ-9.
  • EPI-OBJ-NOTES → OBJ-10.
  • EPI-RTE-PRODUCT-CHECK → RTE-5.
  • EPI-RTE-PRODUCT-VERIFY → RTE-5.
  • EPI-RTE-REVIEW-CREATE → RTE-6.
  • EPI-RTE-REVIEW-SHAPE → RTE-6.
  • EPI-RTE-ADVICE-USE → RTE-7.
  • EPI-RTE-RULE-GENERATE → RTE-8.
  • EPI-RTE-RULE-CHECK → RTE-8.
  • EPI-RTE-RULE-ADOPT → RTE-8.
  • EPI-RTE-RULE-RETAIN → RTE-8.
  • EPI-RTE-RULE-USE → RTE-9.
  • EPI-RTE-EXPERIMENT-CREATE → RTE-10.
  • EPI-RTE-EXPERIMENT-MEASURE → RTE-10.
  • EPI-RTE-EXPERIMENT-DECIDE → RTE-11.
  • EPI-RTE-EXPERIMENT-RETAIN → RTE-11.
  • EPI-RTE-EXPERIMENT-APPLY → RTE-11.
  • EPI-RTE-CONFIDENCE → RTE-12.
  • EPI-RTE-NOTES-WRITE → RTE-13.
  • EPI-RTE-EXPERIMENT-REUSE → RTE-13.
  • EPI-CLM-REVIEW → CLM-3.
  • EPI-CLM-TTSR → CLM-4.
  • EPI-CLM-EXPERIMENT → CLM-5.

Epistemic generation rows marked doctrine-only describe the specialist’s inspected generation boundary. Parent RTE-1/RTE-3 adds an afforded execution route, not observed adherence to review/verification doctrine. Generic product identity merged into OBJ-1; RTE-5 adds diagnostics. Rule functions merge under RTE-8, live use under RTE-9; experiment execution under RTE-10, disposition/retention/apply under RTE-11, note reuse under RTE-13. The sparse ledger retains each function and its weaker evidence separately. Managed skill OBJ-11/RTE-14 was added by parent source inspection as a capability/instruction revision route. No substantive source conflict remains; no model judgment was upgraded to semantic enforcement. The parent returned the narrow matcher-applicability interpretation to the epistemic specialist. Its retained reconciliation confirms conditions 1–4 as wired for that formal subloop, while preserving no observed execution, no corrective-body efficacy and no whole-runtime membership inference. The final synthesis adopts that resolution; the rule policy lifecycle disposition does not deny this narrower formal-solution process.

The specialist supplied no prior classifications to correct. Excluded auxiliary instruction/experiment routes are described by the parent and epistemic pass without expanding the bounded memory profile. Local learn lesson storage belongs to RTE-17/RTE-18/RTE-19; its optional managed-skill branch is separate RTE-14. Generated local consolidation skill files remain OBJ-14 memory-file pull material, not presumed managed-skill installation. Opaque remote checkpoints, raster payloads, remote algorithms and advanced uninspected branches preserve the report’s uncertainty. Filesystem rollout scanning is not generalized to Redis/SQL session storage. Faithfulness absence is limited to this commissioned evidence. Storage and manual-authoring quote coverage was strengthened by the specialist before integration; no assessment was strengthened by the parent.

The runtime/epistemic and memory passes shared a frozen source boundary, not prior-review conclusions. No independent convergence claim is made: overlapping mechanisms are reconciled explicitly. Parent owns generic records and shared admission boundaries; specialists own their respective source interpretations.

Bounded synthesis

oh-my-pi is an enclosing coding-agent runtime whose model calls sit inside substantial symbolic control: context reconstruction, tool interception, approval, steering, task execution and recovery. Its shipped IDE/tool surface supports CLM-1, while CLM-2's performance numbers remain attributed claims. For a long-running coding task, the consequential mechanisms are the separation of scheduler from model, persistent session/checkpoint choices, explicit backend authority differences and worker apply-back. No single approval label describes all alternate effect paths.

For continuity across work, RTE-15, RTE-16, RTE-17, RTE-18, RTE-19 and RTE-20 wire materially different kinds of retention and delivery. The model can receive old traces, generated prose, rasterized context, opaque provider state, project heuristics, ranked facts or prescriptive decisions. The backend comparison profile covers only its named session/backend scope. Its unknown representation, lineage, curation, selector and trace-horizon aggregates preserve uncertainty across included alternatives; exclusions are not silently counted as missing values. Stored lessons can be omitted by budget; hand-edited generated memory can be overwritten; a service clear may leave upstream material intact. These are operationally different failure modes.

For revision, the strongest content-directed loop is narrow: /omfg repairs a formal rule's ability to match historical assistant output. RTE-8 consumes matcher semantics, states a failure when the matcher misses, passes failed content and feedback to the next proposal, and repeats. This supports architecture-level theory-builder membership for matcher-applicability repair, without establishing that it detected the complained-of error or that the corrective body works. Autoresearch provides a different strength: explicit benchmark execution, retained results, agent-decided keep/discard, repository changes and recurring context. A kept variant is an operational choice, not automatically a tested explanation or a harness self-improvement.

Theory-builder conditions 1–4 and learning

The mapping uses theory builder. Each cell gives its own conclusion status. These are route-relative architectural findings, not evidence an actual candidate traversed the route.

Route and proposed solution Condition 1: localized content Condition 2: consumption through content Condition 3: content-directed criticism and resulting change Condition 4: kept result shapes next round Learning
RTE-8, formal matcher/scope as solution to historical applicability wired — identifiable condition/scope plus previous candidate wired — matcher semantics decide which historical surfaces match wired — stated nonmatch/scope failure challenges that candidate; regeneration can replace it wired — failure text and previous rule feed the next proposal/test, across rounds of one generation run uninspected — no candidate-linked before/after capacity evidence
RTE-5, RTE-6 and RTE-7, code adequacy and bug claims afforded — code and report units can state candidate solutions afforded — tools and model context can guide revision, actual content uptake unobserved claimed — review/reproduction doctrine directs substantive criticism; diagnostic acquisition itself is wired afforded — feedback can enter another model turn; actual formulated criticism and uptake not observed uninspected
RTE-10, RTE-11 and RTE-13, experiment variant/playbook hypothesis wired for formal code variants; optional explanation formulation is afforded wired for benchmark execution of variant; use of prose hypothesis is afforded afforded — measurements and notes can support criticism, but status/score alone does not establish it afforded — retained notes/results and kept code support recurrence; no inspected instance establishes retained formulated criticism uninspected — no executed contrast or attribution in this analysis
RTE-14, reusable managed procedure wired — a named, revisable procedure is localized afforded — discovery exposes it for later selection; body use not established here uninspected — structural admission does not criticize efficacy uninspected — replacement/storage alone does not trace criticism into next use uninspected
RTE-16, RTE-17, RTE-18, RTE-19 and RTE-20, retained decisions/guidance afforded — readable parts can carry solutions; opaque branches uninspected afforded — delivered guidance can be used, but delivery alone is not activation uninspected — transformation, source-substring guard or model curation is not by itself formulated criticism of a consumed theory uninspected — later delivery does not establish that a criticism's result shaped the next round uninspected

Addressability is separate: rule condition/scope/body, named skill text, file-level code variants, structured review findings and individual memory entries expose parts for inspection/replacement. RTE-8's symbolic matcher offers finer inspectability than an opaque checkpoint. It does not follow that every retained artifact expresses an explanatory theory. Historical adoption rationale is not required for criticism; where retained, summary decision reasons and sharpshooter delta rationale have explicit later reading routes, while experiment-note rationale depends on supplied content.

Persistence is also separate. Failed-rule feedback reaches subsequent attempts within a generation run; saved rules and managed skills can span sessions. Experiment code, results and notes persist across experiment rounds/resumption. Session/checkpoint and project/bank memory span their recorded scope, with opaque payloads and mixed task histories preventing a complete task-horizon classification. Deployment availability does not prove a later consumer exercised the content.

Learning conclusion status: uninspected. The supported contribution is a narrow criticism-and-retry architecture plus several durable context/revision routes. No inspected comparison shows improved future capacity attributable to criticizing a consumed theory. The memory profile's trace_learning yes is a different, satisfied write/read-back criterion; it is not this learning conclusion.

Reflection conclusion status: wired, restricted to RTE-8/RTE-9's representation of the runtime's own assistant behavior and rule-mediated control. Conversation changes update the history inspected by the rule-generation/checking process; operations on the resulting rule can alter later matching, interruption and prompting. This is a bounded causal representation of behavior, not a representation of every runtime component. The stronger reflective theory-builder qualifier for method efficacy has conclusion status uninspected: matcher applicability checks do not establish criticism of whether the proposed corrective method improves the work.

Autonomy is role-relative. In the narrow matcher repair subloop, generation, evaluation, formulated failure and regeneration are computational (conclusion status: wired). In the encompassing RTE-8 adoption route, the human chooses what to keep and may amend or override, so a fully computational adoption claim has conclusion status inapplicable to that stated human-including route. Autoresearch proposal, measurement, agent judgment, recordkeeping, recurrence and repository transitions are wired computational roles once activated; quality of blame assignment and a complete theory-criticism process remain uninspected. The enclosing runtime receives no universal autonomous theory-builder designation.

Self-improvement has two distinct dispositions under self-improving system. A standing, complaint-responsive rule-revision pathway into later instruction/control is wired at RTE-8/RTE-9's declared objective of correcting the complained-of behavior. Managed procedures and trace-derived guidance offer additional change paths, with evidence-responsive admission varying by route. Occurrent operative self-improvement has conclusion status uninspected because no later execution establishes causal dependence on an admitted change. Improved outcomes are likewise uninspected. Autoresearch's external project optimization does not by itself change oh-my-pi's organization.

The assessment would change with candidate-linked runs showing complaint, failed matcher, regenerated rule, human adoption and later correct/incorrect triggering; experiments linking explicit hypotheses and criticisms to next-round selection; or recall/skill interventions measuring whether retained content changes later work. Frozen backend/provider evidence could resolve opaque classifications. These are evidence requirements, not recommendations to adopt or rank the product.

Limitations

Limitation Affected IDs Inspected boundary Conclusion prevented Resolving evidence
No candidate-linked execution or controlled comparison SRC-1, SRC-2, CMP-2, CLM-2 and all implemented routes Frozen repository only Activation, actual benefit, causal performance attribution and occurrent self-improvement Retained runs connecting specific content/revisions to later behavior, with suitable comparisons
Remote model implementation and version resolution CMP-2 SDK dispatch and configured identities Exact parameter fixity, private criticism or training changes Provider version/weight identity and accessible operation evidence
External backend and opaque checkpoint internals Memory records named in profile Client interfaces and inspected local implementations Complete representation or curation classification for opaque branches Frozen backend/protocol internals and content-preservation evidence
Deployment and arbitrary extensions CMP-3, CMP-4, RTE-2, RTE-3, RTE-4 Registry wrapper, task settings and extension APIs Global permission/containment guarantee Concrete granted settings, host/process boundary, all effect paths and adversarial checks
Bounded tool/revision-family coverage SRC-1, SRC-2 Material routes selected for runtime and both lenses Exhaustive review of every LSP/debugger/browser/provider/MCP/custom command path Further bounded source passes for the specific omitted mechanism
Source predates the analysis date SRC-1, SRC-2 2026-09-05 commit analysed 2026-09-26 Applicability to later revisions A new run pinned to the desired later revision

Verification and blockers

Semantic verification

Checked run/source/method/input/report identities, source-only construction, all declared canonical IDs and source-native referents. No ID was reassigned after sharing; proposals were mapped by exact tokens. Quote blocks retain source text once on canonical records; lens overlays reference them. The source pin is unchanged.

Profile scope was checked against OBJ-12, OBJ-13, OBJ-14, OBJ-15, OBJ-16 and OBJ-17 and RTE-15, RTE-16, RTE-17, RTE-18, RTE-19 and RTE-20. It includes opaque provider checkpoints and service-client alternatives, and explicitly excludes auxiliary RTE-8, RTE-9, RTE-10, RTE-11, RTE-12, RTE-13 and RTE-14 from aggregate memory classification. These exclusions do not remove them from whole-system analysis. Readable summaries were not substituted for operative opaque payloads. SQL/Redis alternatives set the weaker afforded storage basis; local rollout scanning was not assumed to consume those stores.

Trace-fed writes checked: RTE-16 prose/branch summaries, remote replacement history, snapcompact and shake; RTE-17 staged project extraction/consolidation and lesson read-back; RTE-18 retention/fact extraction/sleep transformations; RTE-20 event-fed decision extraction/consolidation. RTE-19 retains service inputs, but upstream derivation/timing remains unknown. The known trace_learning yes follows concrete qualifying routes; aggregate scope/timing/form remain not-determinable across the included opaque and multi-task alternatives. Every dependent axis preserves that boundary. Ordinary raw logging was not counted as trace learning by itself.

Push audit: RTE-15 selects actual branch entries by leaf/parent identities; RTE-16 chooses checkpoint and kept tail. RTE-17 selects current-project summary/lessons automatically under a shared token budget, while detailed memory files are requested pulls. RTE-18 selects bank-scoped relevance using lexical/vector and metadata inputs; RTE-19 passes query/history/tags/budget to an opaque service selector. RTE-20 supplies populated project files coarsely; its requested lexical search is not lexical push. Consumer, trigger, selector input and selected retained parts are explicit. The signal aggregate remains uncertain rather than omitting the opaque alternative.

Operational versus epistemic admission was checked on RTE-2, RTE-3, RTE-5, RTE-6, RTE-7, RTE-8, RTE-11, RTE-14 and RTE-20. Model severity, successful exit, parseable output, Git applicability, matching history and source-substring evidence do not independently establish task correctness or efficacy. Theory-builder conditions, memory trace learning, learning, reflection, autonomy and dispositional/occurrent self-improvement remain separate. No test-source or stored metric was upgraded to an observed run or causal result. Memory authority and rationale differences remain preserved. Explicit uncertainties are analytical limitations, not publication blockers.

Deterministic validation

Validation target: commonplace-validate --full kb/reports/state/agentic-system-analysis/AAS-2026-09-26-oh-my-pi-01/result.md. Structural validation: PASS, clean. Publication also verifies every source anchor and quoted passage against the frozen source. Structural checks do not certify semantic judgments.

Blockers

None.