oh-my-pi agentic-system analysis
Type: kb/types/agentic-system-analysis-result.md
Run identity
Run state: kb/reports/state/agentic-system-analysis/AAS-2026-09-05-oh-my-pi-02/run-state.md
Generated review: kb/agentic-systems/reviews/oh-my-pi.md
Memory analysis report: kb/reports/state/agentic-system-analysis/AAS-2026-09-05-oh-my-pi-02/memory-report.md
Memory analysis report SHA-256: e9219f6db72b36786371ae22a1700db11fb6d6ca3c8d917dd998bfc92f34c6a3
Analysis content is complete when validated; publication completion is declared only by the run state. Coordinator: GPT-6; method: kb/instructions/analyse-agentic-system/SKILL.md and its mandatory epistemic procedure. No prior review prose, prior exact result or substantive audit findings supplied this analysis. Destination eligibility was checked by the publication command. No dynamic system run supplies evidence.
Boundary and evidence
Evidence basis: implementation and shipped documentation at Git commit be6cb8217cd4c1dafcc86793ae5d809ea4d7396a, inspected 2026-09-05. The commit was available locally and dated 2026-09-05. Fetch attempts did not return; applicability beyond this commit is unverified.
The intended use is to explain how the shipped coding agent performs work, admits effects and revisions, and reuses context. The target is an enclosing runtime: CLI/SDK and agent core own the model/tool loop, session state, user steering, extensions, delegated workers and built-in memory integrations. Boundary kind is whole-system, referring to this coding runtime rather than every product in the monorepo. The implementation tier is code-grounded; advertised outcome improvements remain claims.
Included route families: ordinary prompt/tool completion, tool approval and native bridge grants, optional isolated task execution and merge admission, advisor feedback, TTSR interruption, Hashline admission/recovery, extension loading, API retry, built-in autoresearch measurement/disposition, and the integrated memory boundary below. Bun, native Rust bindings, filesystem/process permissions, Git, configured model providers and optional memory services are external contracts. A tool capability is not evidence of a present grant or a deployed isolation envelope. The mapping to behavioral authority names each consumer, channel, force and horizon; it does not replace the source-native rule, tool or memory mechanism.
Excluded: separately deployed python/robomp, metaharness/evaluation campaigns, collaboration/browser relay services, external MCP implementations and external provider/memory-server internals. These exclusions prevent claims about those deployments, evaluator validity, remote storage or model weights. ACP/client controls are inspected at the bridge contract, not in an actual editor. Browser/desktop, LSP/DAP, web acquisition, plan/prewalk, commit automation and security workflow implementations are not exhaustively traced; they remain alternate capabilities and prevent an exhaustive whole-product guarantee. No universal containment, correctness, knowledge-production or performance conclusion is made.
Source register
| Source ID | Kind | Identity/location | Revision | Evidence layer | Inspected scope | Citation anchors | Access gaps and conclusion prevented |
|---|---|---|---|---|---|---|---|
| SRC-1 | Git | https://github.com/can1357/oh-my-pi; operational root /home/zby/llm/commonplace/related-systems/can1357--oh-my-pi |
be6cb8217cd4c1dafcc86793ae5d809ea4d7396a |
Implementation | Selected ranges in packages/agent/src, packages/ai/src, packages/coding-agent/src and crates/pi-edit/src; memory paths enumerated by adopted records | Agent loop, tool gate; every local anchor below resolves at this revision | No worktree evidence, provider internals, deployment traces or causal interventions. Wiring does not establish successful operation. |
| SRC-2 | Git | https://github.com/can1357/oh-my-pi |
be6cb8217cd4c1dafcc86793ae5d809ea4d7396a |
Doctrine/design; promotional outcome claims identified separately | README.md, docs/approval-mode.md, docs/extension-loading.md, docs/ttsr-injection-lifecycle.md, docs/non-compaction-retry-policy.md, memory documentation and shipped prompts | Claims, approval contract | External blog, video captures, benchmark execution and current website not inspected; reported benefits remain claimed. |
No observed-run or causal-experiment source was admitted. Repository listings and searches selected source ranges; truncated tool output was not used as evidence, and relevant content was read again in bounded ranges. Source identities are shared with the fresh memory specialist.
Shared records
Components
| ID | Source-native component | Form/substrate and role | Fixity and evidence |
|---|---|---|---|
| CMP-1 | AgentSession, Agent and agent loop | Symbolic TypeScript in repository; in-process state and orchestration | Implementation conclusion status: wired. Session construction and prompt handoff: SRC-1 packages/coding-agent/src/sdk.ts:3531-3594,3716-3745; packages/coding-agent/src/session/agent-session.ts:6303-6350,6385-6428. |
| CMP-2 | Selected primary/subagent model | Distributed-parametric inference behind configured provider/API; main reasoning, tool choice and task output | Endpoint/configuration resolution conclusion status: wired. Provider/model identity, not immutable weight bytes, is passed per call. Parameter changes during operation: uninspected; provider internals are outside the source. SRC-1 packages/agent/src/agent-loop.ts:1593-1639; packages/coding-agent/src/session/settings-stream-fn.ts:25-89; packages/ai/src/stream.ts:1435-1460. Neither unchanged parameters nor exact weight pinning is established. |
| CMP-3 | Advisor model | Distributed-parametric reviewer with its own agent context | Invocation and delivery conclusion status: wired. Same endpoint-vs-weight limitation as CMP-2; parameter changes and exact weight identity: uninspected. SRC-1 packages/coding-agent/src/sdk.ts:3704-3723; packages/coding-agent/src/session/agent-session.ts:1411-1421; packages/coding-agent/src/session/session-advisors.ts:1185-1288. This is a model judgment, not an answer oracle. |
| CMP-4 | pi-edit Hashline patcher | Symbolic Rust; source snapshots and file-hash recovery | Implementation conclusion status: wired. SRC-1 crates/pi-edit/src/modes/hashline/patcher.rs:176-250. Checks edit anchoring; does not judge program correctness. |
| CMP-5 | ExtensionRuntime / ExtensionToolWrapper | Symbolic JS/TS factories and tool wrappers in process | Loading, registration rollback and per-call gates: wired. SRC-1 packages/coding-agent/src/extensibility/extensions/loader.ts:358-478; packages/coding-agent/src/extensibility/extensions/wrapper.ts:178-335. Runtime-loaded extension behavior beyond the inspected built-ins is uninspected. |
CMP-6 — Auto-thinking classifier. Distributed-parametric component behind an online tiny/smol role or configured local model key. Invocation resolution conclusion status: wired; exact deployed weight pinning and parameter changes during operation: uninspected. Its inference output controls reasoning effort rather than an answer's truth. SRC-1 packages/coding-agent/src/auto-thinking/classifier.ts:115-206; packages/coding-agent/src/session/agent-session.ts:6358-6370.
The memory component roles below share the same source pin. Invocation wiring is established; actual deployment/model weight identity and upstream parameter changes are uninspected. Role selectors are symbolic configuration, not retained weight learning.
CMP-20 — Local memory extraction and consolidation model roles, used by RTE-23. Phase one requests the default role and phase two smol; resolution uses configured role/default selectors, then the active session model or first registry entry. The caller resolves credentials and invokes completion with the static extraction/consolidation prompts. This pins the calling code and role logic, not the model weights or a deployment's actual chosen model. No parameter-update operation is shown in these callers. SRC-1 packages/coding-agent/src/memories/index.ts:359-375,517-532,565-570,724-769,1241-1255; SRC-2 docs/memory.md:98-107.
CMP-21 — Sharpshooter extraction/consolidation model, used by RTE-28. An explicit sharpshooter.model selector resolves through the registry; otherwise the smol role is selected. Both extraction and consolidation use this resolved completion model with their distinct prompts/tool schemas; consolidation records the model ID in its result metadata. Model output supplies semantic admission judgment, while code supplies substring/type checks. Role/model ID and invocation are known; actual parameter bytes, weight revision and provider-side changes are unknown. No weight-training operation is established by extraction or consolidation. SRC-1 packages/coding-agent/src/sharpshooter/extract.ts:151-168,201-215,234-310; packages/coding-agent/src/sharpshooter/consolidate.ts:139-172,183-192.
CMP-22 — Mnemopi generation backend, used for the extraction branch of RTE-25 and any configured model-backed memory work. llmMode: none disables it. A recognized providers.memoryModel local key overrides the online generation branch and is passed to tinyModelClient.complete; otherwise remote supplies configured base URL/model/credentials, or the wrapper resolves tiny then smol and builds a completion callback with current credentials. Failure to resolve a model/key can leave generation unavailable. These are invocation/key/endpoint identities, not verified local weight files or hosted weight pins. The inspected sleep promotion path separately sets llm_used: false; the presence of a consolidation prompt does not make that path parametric. No parameter training is shown by these caller paths; inaccessible underlying parameter changes remain unknown. SRC-1 packages/coding-agent/src/mnemopi/backend.ts:477-484,507-605; packages/mnemopi/src/core/beam/store.ts:306-338; packages/mnemopi/src/core/beam/consolidate.ts:1010-1034.
CMP-23 — Mnemopi embedding model dependency for OBJ-35/RTE-25. The wrapper selects an explicit embedding model, environment override, or variant-derived BAAI/bge-base-en-v1.5 / intfloat/multilingual-e5-large name; configuration also supplies an optional embedding endpoint/key and a no-embeddings switch. These model labels and API/local routes establish embedding invocation configuration, not exact weight bytes or an observed loaded model. Retained vectors are generated access data; rebuilding them after changing the model is not training its weights. Weight pinning and changes inside the embedding implementation/service are not determined here. SRC-1 packages/coding-agent/src/mnemopi/config.ts:53-64,84-95; packages/coding-agent/src/mnemopi/backend.ts:507-520; packages/mnemopi/src/core/beam/store.ts:526-546; SRC-2 docs/mnemosyne-memory-backend.md:61-68,87-91.
CMP-24 — Coding-model memory author and private capture invocation, used by RTE-22, RTE-24, RTE-29 and RTE-31. The private autolearn runner takes the active source agent model, clones current messages/system prompt, allocates a distinct session/provider-state context and invokes that model with capture tools. The ordinary coding model can author rewind findings, lessons and autoresearch notes through their tools. A separate capture session is not a separate verified set of weights or a training run. Actual active-model identity and provider parameters are unobserved; these code paths alter durable artifacts and prompts, without an inspected weight-update call. SRC-1 packages/coding-agent/src/sdk.ts:1204-1255,4205-4212; packages/coding-agent/src/session/agent-session.ts:8044-8091; packages/coding-agent/src/autoresearch/tools/update-notes.ts:24-55.
CMP-25 — Local continuation summarizer, used by the model-generated branches of RTE-21/RTE-22. The summary function receives a resolved Model and credential, serializes history, accounts for its window and calls the summary model over bounded windows, carrying an earlier summary forward when necessary. Split-turn history/prefix summaries are then combined. These are model-input/output transformations whose durable result is OBJ-22/OBJ-25, OBJ-26; model parameters are not modified by the inspected summarizer. The precise runtime model/weight revision and its upstream updates are unobserved. Snapcompact and shake are mechanical alternatives, not additional summarizer models. SRC-1 packages/agent/src/compaction/compaction.ts:854-902,1791-1836; SRC-2 docs/compaction.md:3-8,136-150.
CMP-26 — Remote compaction provider dependency, used by RTE-21 to produce OBJ-24. The runtime selects a supported provider-native/configured remote compaction route and retains its returned replacement history for provider replay. The adapter-visible provider/model/endpoint route is an invocation identity. The encrypted state is neither an inspected weight file nor evidence of parameter training; its provider-side encoding, model architecture, exact weights and any parameter changes are inaccessible. Remote success skips local generation of a second substantive summary. SRC-1 packages/coding-agent/src/session/compaction-methods.ts:11-15,102-105; packages/agent/src/compaction/openai.ts:1-12,884-907; packages/agent/src/compaction/compaction.ts:1794-1802; packages/coding-agent/src/session/session-context.ts:165-177,424-477.
CMP-27 — Hindsight service-side memory reasoning dependency behind RTE-26/RTE-27. The supplied adapter calls bank-scoped retain/recall and mental-model create/list/refresh interfaces, with service configuration and bounded query/render inputs. It does not expose which parametric models, embeddings or nonparametric algorithms the server uses for extraction, reflection or selection. The API invocation is wired; internal component composition, exact model/weight identity and parameter-change behavior are uninspected. Register this as an opaque external dependency rather than inventing a specific model. SRC-1 packages/coding-agent/src/hindsight/state.ts:293-307,326-371,459-491; packages/coding-agent/src/hindsight/mental-models.ts:142-175,215-243; SRC-2 docs/memory.md:130-152.
Operative objects
| ID | Source-native part | Form / storage / lineage | Producer → consumer, authority and evidence |
|---|---|---|---|
| OBJ-1 | Assistant answer/proposed explanation | Natural-language within structured message; active context and session persistence; model-generated from prompt/history | CMP-2 → user/later model. Advisory truth-apt content, not warranted by completion. SRC-1 packages/agent/src/agent-loop.ts:1438-1555; packages/coding-agent/src/modes/print-mode.ts:154-180,187-231. |
| OBJ-2 | Tool-call name and arguments | Symbolic JSON plus possible natural-language content; active message; generated by model or host bridge | Caller → tool validator/wrapper. Requests an effect, not a grant by itself. SRC-1 packages/coding-agent/src/extensibility/extensions/wrapper.ts:178-267. |
| OBJ-3 | Tool result / command output | Text and structured status, potentially mixed/opaque payload; active context then session records | Executor/provider → model. Evidence about the named invocation only, with external-source warrant uninspected. SRC-1 packages/agent/src/agent-loop.ts:1438-1453,2586-2619. |
| OBJ-4 | Proposed and changed workspace content | Symbolic code/configuration and natural-language files, on disk or optional isolated worktree/Git branch | Model/tool → filesystem, tests, parent/operator. Direct executable or advisory force depends on the file consumer. SRC-1 packages/coding-agent/src/task/structured-subagent.ts:598-672; crates/pi-edit/src/modes/hashline/patcher.rs:176-250. |
| OBJ-5 | Subagent result and structured-output status | Structured JSON/text plus patch/branch references; assignment artifact storage | Worker → parent. Schema status and exit status are distinct from substantive correctness. SRC-1 packages/coding-agent/src/task/structured-subagent.ts:290-335,629-681. |
| OBJ-6 | Advisor note and severity | Natural-language note plus symbolic severity; aside/steering/custom-message channel | CMP-3 → primary model and UI. May interrupt or resume scheduling; content remains advice. SRC-1 packages/coding-agent/src/session/session-advisors.ts:1185-1288. |
| OBJ-7 | Extension module / registrations | Symbolic JS/TS repository or configured files; authored externally or locally | Module factory → runtime tools/providers/hooks. Operational authority upon loading; semantic trust not established by successful import. SRC-1 packages/coding-agent/src/extensibility/extensions/loader.ts:358-478. |
| OBJ-8 | Experiment run output and parsed metrics | Natural-language command output and symbolic numeric fields; logs/storage; acquired from benchmark execution | Benchmark → model and logger. Measures the chosen harness; meaning and validity depend on that harness. SRC-1 packages/coding-agent/src/autoresearch/tools/run-experiment.ts:57-90,150-209. |
| OBJ-9 | Logged experiment status, metric and scope deviations | Structured retained experiment record; selected by calling model; linked to run/commit | Caller/logger → future prompt, Git admission and metric summaries. Logged metric may differ from parsed metric. SRC-1 packages/coding-agent/src/autoresearch/tools/log-experiment.ts:64-200; SRC-2 packages/coding-agent/src/autoresearch/prompt.md:20-45,53-96. |
| OBJ-10 | TTSR rule body, match condition and injection | Natural-language rule plus symbolic regex/AST conditions; loaded file/manager, retained injection | Author/stream selector → primary model and abort coordinator. Conditional instruction and scheduling force; no truth endorsement. SRC-1 packages/coding-agent/src/session/ttsr-coordinator.ts:389-465; SRC-2 docs/ttsr-injection-lifecycle.md:19-76. |
The following memory objects split content, access structures and opaque payloads where their producers or consumer authority differ. Storage adapters are not evidence of a deployed service; source-native labels are preserved. Unless stated otherwise, implementation conclusion status is wired and operation conclusion status is uninspected.
OBJ-20 — Retained session messages: raw user/assistant/tool/custom content, with linguistic and potentially opaque/image payloads; acquired from use or imported foreign sessions. Ordinary file store plus in-memory and afforded SQL/Redis SessionManager adapters. Later coding-model replay and extraction consumers; advisory/learning input, not truth endorsement. SRC-1 packages/coding-agent/src/session/session-storage.ts:15-43,58-91,103-151; packages/coding-agent/src/session/session-manager.ts:2795-2818,3105-3127; packages/coding-agent/src/session/foreign-session-import.ts:35-51.
OBJ-21 — Session parent/entry IDs, reset/compaction boundaries and import provenance: symbolic access metadata in the same storage interface, routing history selection. SQL/Redis choices are afforded for the named SessionManager consumer. SRC-1 packages/coding-agent/src/session/sql-session-storage.ts:8-54,245-254; packages/coding-agent/src/session/redis-session-storage.ts:15-43,94-108; packages/coding-agent/src/session/session-context.ts:180-235,280-305.
OBJ-22 — Linguistic continuation summary/short summary: trace-extracted model text in retained compaction entry, consumed in later coding-model calls; prior summary may feed the next summary. Its first-kept/replay access markers belong to OBJ-21. SRC-1 packages/agent/src/compaction/compaction.ts:1791-1836; packages/coding-agent/src/session/session-maintenance.ts:1492-1537.
OBJ-23 — Snapcompact archive: mechanically transformed transcript text and PNG frames carrying linguistic content, with symbolic archive shape/budget data; retained under preserveData and reattached after reconstruction. A visible lead-in is not the full archive. SRC-1 packages/coding-agent/src/session/session-context.ts:156-162,424-443; SRC-2 docs/compaction.md:142-151.
OBJ-24 — Remote continuation payload: provider replacement-history items retained under preserveData, replayed to the provider. Partly checked structure surrounds opaque encrypted state; complete representational form is not determinable. Remote success bypasses a local summary-model call. SRC-1 packages/agent/src/compaction/openai.ts:1-12; packages/agent/src/compaction/compaction.ts:1794-1802; packages/coding-agent/src/session/session-context.ts:165-177,424-477.
OBJ-25 — Branch navigation summary: model-generated linguistic description of abandoned context retained at the new branch, consumed during reconstruction. SRC-1 packages/coding-agent/src/session/agent-session.ts:9569-9599; packages/coding-agent/src/session/session-context.ts:371-373.
OBJ-26 — Rewind findings: investigator-model report from tool investigation, retained as branch summary/custom message after rewind. Guidance for the continuing model; not established factual acceptance. SRC-1 packages/coding-agent/src/session/agent-session.ts:8044-8091.
OBJ-27 — Local MEMORY.md and memory_summary.md: trace-extracted/consolidated project language files. Summary is startup guidance; longer document can be requested. SRC-1 packages/coding-agent/src/memories/index.ts:182-226,957-1002,1280-1297.
OBJ-28 — Local learned.md: explicit learned lessons plus preserved literal human edits; linguistic file, authored or trace-derived. Newest-first bounded bullets and retained non-bullet material feed a later startup snapshot. SRC-1 packages/coding-agent/src/memories/index.ts:1284-1297,1367-1419.
OBJ-29 — Local generated skills and optional assets: SKILL.md language plus potentially symbolic scripts/templates/examples in project memory files, written by consolidation for later requested use. SRC-1 packages/coding-agent/src/memories/index.ts:957-1002; SRC-2 packages/coding-agent/src/prompts/memories/consolidation.md:7-30.
OBJ-30 — Local phase-one extraction outputs: raw-memory text (already a model extraction), synopsis and optional slug stored in SQLite; consumed by consolidation as learning input. SRC-1 packages/coding-agent/src/memories/index.ts:345-410,724-807; packages/coding-agent/src/memories/storage.ts:48-72.
OBJ-31 — Local thread/job/lease state: symbolic SQLite provenance and eligibility metadata governing extraction/maintenance, separate from the extracted claims. SRC-1 packages/coding-agent/src/memories/storage.ts:48-72; packages/coding-agent/src/memories/index.ts:345-410,724-807.
OBJ-32 — Mnemopi working transcripts: retained linguistic multi-author text in SQLite with source/session/cwd metadata and initial unknown veracity. User-only extraction and marker-free embedding projections are distinct inputs. SRC-1 packages/coding-agent/src/mnemopi/state.ts:522-551; packages/mnemopi/src/core/beam/store.ts:306-338.
OBJ-33 — Mnemopi episodic summaries: bounded mechanically combined linguistic content, SQLite persistence, source IDs and aggregated veracity. Working-to-episodic promotion is not itself new knowledge. SRC-1 packages/mnemopi/src/core/beam/consolidate.ts:1010-1048.
OBJ-34 — Mnemopi extracted facts/preferences/instructions/timelines/triples: mixed linguistic assertions and symbolic structures in SQLite, produced by extraction, later recalled as background context. A fact category named instruction does not grant normative priority over the current user. SRC-1 packages/mnemopi/src/core/beam/consolidate.ts:389-430,524-550; packages/coding-agent/src/mnemopi/backend.ts:138-150.
OBJ-35 — Mnemopi access structures: full-text indexes, vectors, episodic/entity/knowledge graph links, importance/recency/expiry/scope/source IDs and in-memory query/access state. Symbolic routing/ranking structures; numerical embeddings are not updated model weights. SRC-1 packages/mnemopi/src/core/beam/store.ts:267-271,436-449,526-547; packages/mnemopi/src/core/beam/recall.ts:540-573,617-641,732-756,915-929,1049-1071.
OBJ-36 — Hindsight retained documents/returned facts: adapter-visible remote service objects, imported/exported language and structured response data. Internal server content form, transformation and ranking remain outside evidence. SRC-1 packages/coding-agent/src/hindsight/state.ts:293-307,326-371,425-455.
OBJ-37 — Hindsight evolving mental-model content: returned language strings from remote bank objects, supplied as background knowledge; authored seed configuration is separate static input. Server derivation remains opaque. SRC-1 packages/coding-agent/src/hindsight/mental-models.ts:89-105,142-175,215-250.
OBJ-38 — Hindsight bank/tag IDs and local snippet/model cache: symbolic selection/access data, memory and service-interface metadata, distinct from returned claim content. SRC-1 packages/coding-agent/src/hindsight/state.ts:293-307,425-491; packages/coding-agent/src/hindsight/mental-models.ts:229-243.
OBJ-39 — Sharpshooter queued decision delta: extracted statement, literal supporting user-prompt substring, source/kind/friction fields, optional rationale/alternative and session/time provenance in files. Typed provenance supports occurrence, not entailment. SRC-1 packages/coding-agent/src/sharpshooter/extract.ts:234-310.
OBJ-40 — Sharpshooter architecture.md/product.md/style.md: project language files rewritten from old decisions, queued deltas and project documents; later loaded as instructions unless the user overrides. Successful consolidation consumes delta files, and final bullets need not preserve literal source IDs/quotes. SRC-1 packages/coding-agent/src/sharpshooter/consolidate.ts:123-194; packages/coding-agent/src/sharpshooter/backend.ts:117-129; SRC-2 packages/coding-agent/src/prompts/memories/sharpshooter-consolidate-system.md:1-43.
OBJ-41 — Managed learned skill body: global agent-directory language files, separate from authored skills, created/updated/deleted by capture tools. Body becomes procedural instruction after selection; writer checks size and linked-path safety. SRC-1 packages/coding-agent/src/autolearn/managed-skills.ts:1-26,112-174,178-250; packages/coding-agent/src/tools/manage-skill.ts:55-98.
OBJ-42 — Managed skill name/description catalog: symbolic addressing plus linguistic selection description, discovered automatically; authored same-name skill wins. Availability is separate from selected body delivery. SRC-1 packages/coding-agent/src/extensibility/skills.ts:49-81,344-384,497-523.
OBJ-43 — Autoresearch durable notes/playbook/ideas: model-authored linguistic SQLite content replaced/appended by update_notes and supplied on subsequent iterations. Experiment data remain OBJ-8/OBJ-9; their access state is OBJ-44. SRC-1 packages/coding-agent/src/autoresearch/storage.ts:199-264; packages/coding-agent/src/autoresearch/tools/update-notes.ts:24-55; packages/coding-agent/src/autoresearch/index.ts:299-355,378-414.
OBJ-44 — Autoresearch branch/session/segment/run selection state: symbolic SQLite fields selecting notes, current-segment recent/best/baseline/violations/pending data for the next prompt. SRC-1 packages/coding-agent/src/autoresearch/storage.ts:199-264; packages/coding-agent/src/autoresearch/index.ts:299-355,378-414.
OBJ-45 — Fired-rule names/history: symbolic session access history restored to the TTSR manager on startup. Rule/body/injection remains OBJ-10; no new rule induction is established. SRC-1 packages/coding-agent/src/session/session-context.ts:280-293; packages/coding-agent/src/sdk.ts:1678-1679.
OBJ-46 — Approved-plan path/body: retained local plan language and recovery reference, produced by prior planning and reloaded as instructions after history rewrite. The sent flag prevents repeated delivery until reset; plan existence alone does not establish trace learning. SRC-1 packages/coding-agent/src/session/agent-session.ts:5666-5691,6246-6249,6404-6416; packages/coding-agent/src/session/session-maintenance.ts:1524-1527.
Routes
All runtime routes below have implementation conclusion status wired unless a field explicitly says otherwise. Operation/activation/benefit conclusion status is uninspected: no execution trace was admitted. Guarantee strength is stated separately. Memory routes retain their specialist-specific evidence basis.
RTE-1 — Prompt → provider → tool/result loop → terminal answer. Trigger/principal: CLI/SDK user or authorized host input. CMP-1 owns next-step scheduling; CMP-2 chooses content/tools probabilistically from system prompt, messages and granted tools. AgentSession waits for memory transition, allows before_agent_start context changes, compacts as needed, and invokes the recovery-owned prompt path. Results re-enter context; steering, follow-ups and asides can extend the loop. Terminal output is answer/events/error, consumed by caller. Persistence/read-back: session routes below; immediate state alone is not memory. Delegated visibility is supplied by RTE-3, not implicit shared model context. Selection/expiry: active context/queued messages; per-turn prompt override cleared at unwind. Recovery: RTE-12; no blanket rollback of external effects. Strength: protocol for loop progression, no correctness guarantee. SRC-1 packages/coding-agent/src/session/agent-session.ts:6303-6428; packages/agent/src/agent-loop.ts:1135-1209,1438-1555,1593-1639; packages/coding-agent/src/modes/print-mode.ts:154-180,187-231.
RTE-2 — Tool-call admission → registered executor. Trigger: OBJ-2 reaches wrapper, including nested/direct bridge dispatch. Owner: wrapper plus extension callbacks and UI; symbolic policy resolves mode, tool declaration and explicit user policy, then checks effective arguments after extension revision. Deny rejects; required prompt without UI rejects; pending provider computer checks require UI even with permissive mode. Executor reaches ambient tool effects after admission. Current grants are session registry/configuration, not all advertised capabilities. Immediate return: result/error to RTE-1; later read-back via retained session, no independent approval-memory claim. Delegation: RTE-3 overrides mode but inherits explicit policies. Expiry: per call except ACP client permission caching; deployed settings uninspected. Revision admission: governs ordinary tool-written code, instructions and configuration; caller proposes, wrapper/user can veto, downstream tool checks may reject, user/model can repair. No general effect rollback. Strength: policy at registered-call boundary; no process/filesystem containment. SRC-1 packages/coding-agent/src/sdk.ts:2795-2812,2918-2948; packages/coding-agent/src/extensibility/extensions/wrapper.ts:178-335; SRC-2 docs/approval-mode.md:14-72,134-162.
RTE-3 — Task assignment → independent worker → parent result. Trigger: parent task/Eval subagent request. Owner: structured-subagent preflight and executor; symbolic discovery/model/schema/isolation policy surrounds a model-driven worker. Worker model role resolved from request, settings, agent definition and parent; worker gets its own session/tools and selected task input. Non-isolated runSubprocess is a material branch. Isolated request must be enabled, then gets worktree execution/capture. Immediate return: OBJ-5 with output/error and recovery artifacts. Later read-back/delegated visibility: parent result/artifact interface; worker-memory routes are separately scoped. Selection: assignment identity and options; persistence: artifact lease and optional patch/branch, not automatic acceptance. Recovery: error/abort status, artifacts where available; expiry/parked worker reuse beyond inspected ranges uninspected. Headless worker mode is yolo while explicit tool policies remain. Strength: protocol for dispatched assignments, not guaranteed isolation or valid conclusions. SRC-1 packages/coding-agent/src/task/structured-subagent.ts:290-335,585-681; packages/coding-agent/src/task/executor.ts:944-985.
RTE-4 — Isolated worker changes → parent workspace. Trigger: completed RTE-3; owner: structured executor, with requested/configured apply flag. Admission requires isolation context, applyChanges, exitCode zero, no error and no abort. The caller/model selects apply/merge policy; code can reject merge failures, and apply=false preserves a patch/branch for later choice. OBJ-5 structured-output validity is computed separately from this merge condition, but the upstream finalizer can turn schema violations into exit code 1, thereby blocking this path. This is structural validation, not substantive correctness acceptance. Immediate return: merge summary and changesApplied; later consumer: parent tools/model and workspace execution. Delegated visibility: worker delta, not general sibling state. Invalidation: merge conflicts/recovery artifacts; external side effects are outside worktree rollback. Strength: conditional integration protocol; no semantic acceptance gate. SRC-1 packages/coding-agent/src/task/structured-subagent.ts:629-681; packages/coding-agent/src/task/executor.ts:666-780.
RTE-5 — Primary turn → advisor judgment. Trigger: primary turn end and configured advisor; next-step owner: advisor runtime/CMP-3, with source context and separate agent state. Output: OBJ-6, proposing concern rather than providing an expected answer. Immediate return goes to RTE-6; no independent learning claim for merely running a second model. Persistence/selection/expiry: advisor session/context and runtime lifecycle, detailed provenance/read-back beyond inspected anchors uninspected. Delegated visibility: project context is explicitly supplied. Rejection: emission filter/dedupe, RTE-6; human can control session/advisor. Strength: best effort critique. SRC-1 packages/coding-agent/src/sdk.ts:3704-3723; packages/coding-agent/src/session/agent-session.ts:1411-1421; packages/coding-agent/src/session/session-advisors.ts:1185-1239.
RTE-6 — Advisor note → scheduling and primary context. Trigger: emitted note passes emission guard; owner: symbolic channel selector, based on severity, streaming/abort/idle/plan/client state and immunity. Nit becomes aside; concern/blocker may steer, preserve a visible card, or trigger a turn. Terminal-answer concern and blocker have different wake behavior. Immediate return: recorded note/card or message; later consumer: primary model through custom/steering message, with semantic compliance uninspected. Delegated visibility: primary sees the note, not necessarily advisor's full derivation. Suppressed note has no delivered force; session transition and per-update filter constrain use. Revision proposed: reconsider behavior; primary decides substantive response, human can interrupt, code governs delivery only. Strength: protocol for scheduling, best effort for correction. SRC-1 packages/coding-agent/src/session/session-advisors.ts:1185-1288.
RTE-7 — Configured/discovered extension → runtime capabilities/context. Trigger: startup/load; proposer: file author/configurer; admission owner: loader validates factory shape and executes it, records per-path errors, binds sequentially. JS/TS factories can register tools/providers/hooks; loaded hooks can revise tool input or turn prompt (RTE-2/RTE-1). Storage: source module files and runtime registries. Immediate return: definitions/errors; later use through calls/hooks, with no learned-memory claim for shipped code. Selection: discovered/configured paths; disable/remove/reload changes selection, deployed use uninspected. Provider registration queue rolls back on factory error; arbitrary module side effects are not proven reversible. Strength: module/registration protocol, not a sandbox or semantic review. SRC-1 packages/coding-agent/src/extensibility/extensions/loader.ts:358-478; packages/coding-agent/src/session/agent-session.ts:6308-6350; SRC-2 docs/extension-loading.md:26-76.
RTE-8 — Autoresearch benchmark → measurement record. Trigger: active experiment and run_experiment; owner: tool executes fixed bash autoresearch.sh, records pre-run dirty paths and run identity, captures command output and parses METRIC/ASI fields. Benchmark author/model supplies evaluator and inputs. Immediate result: measured fields, status, log location and tail to model; storage retains run data and later iterations consume it. Selection: active session for current branch; pending prior run can be abandoned. Execution success means exit zero without kill. Recovery: logs, timeout/error data, RTE-9 disposition; benchmark validity is uninspected. Delegated visibility: none specifically established. Strength: measurement protocol for this command, not an answer oracle or proof of improvement. SRC-1 packages/coding-agent/src/autoresearch/tools/run-experiment.ts:57-90,150-212; SRC-2 packages/coding-agent/src/autoresearch/prompt.md:14-39.
RTE-9 — Experiment judgment → keep/discard and successor state. Trigger: log_experiment with a pending run. Calling model proposes status, metric, description, optional flags/justification; logger decides operational consequences. On dedicated autoresearch branch, keep commits modified files; other statuses revert current uncommitted iteration to HEAD. Off-branch keep leaves edits and discard uses a lossy dirty-path filter. Scope deviations and disagreement between supplied and parsed metric cause warnings, not rejection. No enforced metric-improvement comparison in this admitting path. Immediate return: record/warnings; later read-back: future experiment prompt and best-metric summary, excluding flagged runs. Selection: active branch/session and supplied run IDs; expiry: new segment/flagging/clear protocol. Human sets/stops goal; model diagnoses, proposes and decides keep, code performs commit/reset and can return errors. Strength: operational disposition protocol, model policy for honest improvement; no semantic acceptance guarantee. SRC-1 packages/coding-agent/src/autoresearch/tools/log-experiment.ts:64-200,341-379; SRC-2 packages/coding-agent/src/autoresearch/prompt.md:20-45,53-103.
RTE-10 — TTSR match → interrupted generation and rule injection. Trigger: configured scoped regex/AST condition on monitored output. Owner: symbolic matcher/coordinator; author sets rule, model does not decide the match. Coordinator aborts permitted matches, checks generation/retry token, optionally discards partial answer, appends hidden custom instruction and persists rule names/body, then continues. Immediate return: retry/event; later read-back: retained injections and memory-accounted session replay. Selection: rule conditions/scope/repeat settings, not general semantic error detection; AST only sees reconstructed source-bearing payload described in docs. Disabled/unmatched rules do not activate. Already completed unrelated external effects are not undone by removing partial context. Strength: protocol for matching/interruption; policy/best effort for model compliance. SRC-1 packages/coding-agent/src/session/ttsr-coordinator.ts:389-465; SRC-2 docs/ttsr-injection-lifecycle.md:44-76,89-120.
RTE-11 — Hashline candidate → matching/recovery → patch result. Trigger: edit; owner: Rust patcher, using live file hash, stored snapshot, anchor scope and optional seen-line guard. Matching content applies; stale block resolution can use stored snapshot, head/tail edits may apply with warning, anchor edits may recover, otherwise mismatch rejects. Immediate return: patched text or error/warnings to edit caller. Later use: changed workspace and correction information; store retains snapshots for recovery but this route does not establish learning or correctness. Selection: path/hash/anchor, stored version; expiry outside inspected function uninspected. Human/model proposes edit; deterministic guard can veto, caller can re-read/revise. Strength: constrained edit-admission protocol, neither universal stale-edit rejection nor program validation. SRC-1 crates/pi-edit/src/modes/hashline/patcher.rs:176-250.
RTE-12 — Provider failure → bounded retry/fallback or stop. Trigger: error and session recovery; symbolic classification excludes context overflow, applies replay-unsafety veto unless positive evidence says emitted tools did not execute. Model/credential fallbacks are configuration-driven; no model-weight modification is implied. Immediate return: retry/continuation/error. Persistence: prior transcript/error state; later read-back follows session routes. Selection: error kind and replay state; expiry: budgets/generation/cancellation. Delegated invocation has its own runtime settings. Recovery is itself this route, not global effect rollback. Strength: protocol with external provider/effect-reporting contract; exactly-once effects are not established. SRC-1 packages/coding-agent/src/session/turn-recovery.ts:1223-1273; SRC-2 docs/non-compaction-retry-policy.md:18-89.
RTE-13 — User prompt → optional auto-thinking effort. Implementation conclusion status: wired. Trigger: a real user turn with auto-thinking enabled. CMP-6 classifies preprocessed input using an online role model or local model key; symbolic parsing/clamping constrains the selected effort to the target model's supported ceiling. Consumer: current primary model configuration. Immediate return: effort or classification failure; no retained learned artifact or later read-back is established by this route. Delegated visibility and memory expiry are inapplicable to this per-turn setting. Caller handles failures with a concrete fallback; no weight update/quality improvement established. Strength: bounded setting-selection protocol, not verified optimal effort. SRC-1 packages/coding-agent/src/auto-thinking/classifier.ts:115-206; packages/coding-agent/src/session/agent-session.ts:6358-6370.
The following memory routes have implementation conclusion status wired and operation conclusion status uninspected. Their return, maintenance, visibility and revision-admission audit follows in the memory lens; that audit extends these records.
RTE-20 — Automatic history assembly. On resume/rebuild/next context assembly, SessionManager uses leaf identity and parent links, latest reset/compaction and first-kept/replay-through entry IDs to select OBJ-20, OBJ-21, OBJ-22, OBJ-23, OBJ-24, OBJ-25, OBJ-26 for the coding model. This is identifier-based push, including operator-selected branch inputs. Raw trace replay remains distinct from learned summaries. SRC-1 packages/coding-agent/src/session/session-context.ts:180-235,280-305,404-477.
RTE-21 — Compaction/archival learning and replay. Manual, overflow, length, threshold, mid-turn or idle triggers transform history into OBJ-22, OBJ-23 or OBJ-24 and persist an entry; OBJ-20, OBJ-21 reconstruction then supplies it to later coding-model calls. Soft/split summaries call a model; snapcompact is mechanical archival; remote providers return opaque state. Handoff is an alternative generated continuation document; shake/pruning mechanically shorten retained history and provide recovery references. Each has a later consumer, but source/provider representations differ. SRC-2 docs/compaction.md:59-68,123-150,157-180; SRC-1 packages/agent/src/compaction/compaction.ts:1791-1836; packages/coding-agent/src/session/compaction-methods.ts:11-45; packages/coding-agent/src/session/session-maintenance.ts:1492-1537.
RTE-22 — Navigation/rewind learning. Branch navigation can summarize abandoned context; checkpoint work instead lets the investigator model compose a rewind report from its tool investigation. Both write OBJ-25, OBJ-26, which RTE-20 supplies after the branch transition. Continuation after the investigation is established; a universal session-to-task mapping is not. SRC-1 packages/coding-agent/src/session/agent-session.ts:8044-8091,9569-9599; SRC-2 docs/compaction.md:3-8,27-55.
RTE-23 — Local historical extraction, maintenance and startup supply. Top-level persisted-session startup scans eligible older sessions, skipping current/too recent/too old or over-budget work, and uses a model to write extraction content OBJ-30 while symbolic job/provenance state OBJ-31 controls the work. A second pass writes OBJ-27 and OBJ-29; lessons OBJ-28 have the independent RTE-24 writer. At prompt construction, cwd chooses the project root, availability chooses summary/lessons, then a shared approximate token budget clips them. The coding agent receives Memory Guidance and can request full files. Startup consolidation may refresh the summary snapshot while preserving the earlier lesson snapshot. SRC-1 packages/coding-agent/src/memory-backend/local-backend.ts:19-35; packages/coding-agent/src/memories/index.ts:182-226,251-270,345-410,724-807,957-1002; SRC-2 docs/memory.md:24-28,70-84,111-128.
RTE-24 — Explicit lesson capture and manual editing. A model learn call uses backend save; the local branch normalizes/redacts, bounds fields, removes duplicate bullet text, inserts newest-first and caps 100 lessons. Manual non-bullet edits survive. RTE-23 delivers lessons from the next session snapshot rather than mutating the active prompt prefix. Acquisition is separate from deduplication/oldest-entry forgetting over the existing file. SRC-1 packages/coding-agent/src/memory-backend/local-backend.ts:14-35; packages/coding-agent/src/memories/index.ts:1284-1297,1367-1419; SRC-2 docs/memory.md:57-66.
RTE-25 — Mnemopi retention, extraction, promotion and recall. Agent-end cadence or explicit enqueue sends unretained transcript portions into bank-scoped working rows, with user-only extraction and marker-free embedding projections. Background extraction writes facts/relations; age-eligible sleep promotes working rows to episodic records. First-turn and pre-compaction selectors compose a query from current/recent user turns, search configured banks, merge and cap results, and inject formatted background text into the coding agent or compactor. Explicit tools use the same scoped access. SRC-1 packages/coding-agent/src/mnemopi/state.ts:385-431,466-551,573-620,626-705; packages/mnemopi/src/core/beam/store.ts:306-338,526-547; packages/mnemopi/src/core/beam/consolidate.ts:972-1048; packages/coding-agent/src/mnemopi/backend.ts:138-150,316-318.
RTE-26 — Hindsight retention and query recall. Agent-end cadence or forced enqueue prepares full-session or sliding-window transcripts, includes bank/tags/source timestamp, and requests asynchronous server retention. The first-turn and compaction hooks compose queries and automatically import bounded returned facts. The later consumers are the coding model and compaction model; server transformation and query ranking remain unknown. Explicit retain/recall/reflect are separate model requests. SRC-1 packages/coding-agent/src/hindsight/state.ts:293-307,326-384,425-455; packages/coding-agent/src/hindsight/backend.ts:86-108,134-146; SRC-2 docs/memory.md:144-152.
RTE-27 — Hindsight mental-model supply and maintenance. Startup lists content from the active bank; optional seeding creates missing models only. Local tag matching includes matching tagged and untagged models, then name sorting and a character budget choose rendered content. Prompt rebuild supplies it to the coding model without that model requesting it; agent-end TTL expiry can refresh the snapshot. Explicit reload/refresh/delete/reseed are operator surfaces. Listing/reloading is not itself synthesis; server-side derivation is opaque. SRC-1 packages/coding-agent/src/hindsight/mental-models.ts:142-175,215-243,275-307; packages/coding-agent/src/hindsight/state.ts:459-520; SRC-2 docs/memory.md:44-55.
RTE-28 — Sharpshooter extraction, consolidation and instruction supply. A committed user-message event selects the prompt plus bounded previous-human/assistant context; extraction is skipped for slash commands, short prompts, disposed sessions or an already-running extraction. Admitted deltas persist in OBJ-39. A scheduled/forced model consolidation reads all pending deltas, prior decision files and project docs, replaces OBJ-40 and consumes deltas. Prompt rebuilding reads populated project files and truncates to the configured injection budget. SRC-1 packages/coding-agent/src/sharpshooter/backend.ts:78-125,132-144; packages/coding-agent/src/sharpshooter/extract.ts:98-148,171-198,234-310; packages/coding-agent/src/sharpshooter/consolidate.ts:123-194.
RTE-29 — Autolearn capture and managed-skill return. With settings enabled, a substantive top-level turn can trigger a separate capture model run using a copy of the source agent's messages and system prompt. It receives capture tools that can save lessons or manage skills. The threshold counts tool completions; aborted/plan/goal turns are excluded, and autoContinue must be enabled. Generated skills persist globally for future discovery. Catalog supply is coarse push; reading/using a particular discovered skill is a model request or operator invocation. This is cross-task-capable procedural memory with no demonstrated reuse success. SRC-1 packages/coding-agent/src/autolearn/controller.ts:73-149; packages/coding-agent/src/sdk.ts:1204-1255,4191-4215; packages/coding-agent/src/autolearn/managed-skills.ts:1-26; packages/coding-agent/src/extensibility/skills.ts:18-29,49-81,344-384,497-523.
RTE-30 — Explicit pull interfaces. The coding agent reads local memory files with read memory://root/...; Mnemopi permits a full stored row read by memory ID, as distinct from clipped recall previews. Recall/reflect tools are supported for Mnemopi/Hindsight; local memory exposes neither structured search nor those tools. Agents are instructed to read the full row before replacement updates; no enforced read-before-update proof is established here. SRC-1 packages/coding-agent/src/internal-urls/memory-protocol.ts:16-40,114-129; packages/coding-agent/src/memory-backend/local-backend.ts:34-44; SRC-2 docs/memory.md:30-42,57-66,148-150; docs/mnemosyne-memory-backend.md:31-40,81-87.
RTE-31 — Autoresearch iteration memory. The model replaces notes/appends an idea through update_notes; current branch chooses the active stored research session. Before an iteration the hook reloads logged runs, reconstructs state, selects current-segment last-three/best/baseline/violations/pending data and supplies notes/results in the system prompt. Thus observed experiment feedback can shape a persisted model-authored playbook and later iterations. The store write is wired; whether any particular note incorporates results is unobserved. This concerns builtin experiment mode, not an external evaluation campaign. SRC-1 packages/coding-agent/src/autoresearch/tools/update-notes.ts:24-55; packages/coding-agent/src/autoresearch/index.ts:299-355,378-414; SRC-2 packages/coding-agent/src/autoresearch/prompt.md:14-30,47-76.
RTE-32 — Approved-plan reload. History reset/compaction clears the sent flag. Before a subsequent prompt, the local plan reference path selects the stored plan; when not in plan mode and not yet sent, the runtime inlines its body and path. The flag is committed at prompt delivery. The executor is the named later consumer, making this identifier-based push, while the path also affords requested recovery reads. This route does not prove trace-derived learning. SRC-1 packages/coding-agent/src/session/agent-session.ts:5666-5691,6246-6249,6404-6416; packages/coding-agent/src/session/session-maintenance.ts:1524-1527.
RTE-10 memory amendment — OBJ-45 adds accumulated fired names to the existing rule/injection route: reconstruction gathers names and startup restores them to the TTSR manager. This changes future rule selection, without inducing rules. The earlier record described injection persistence alone; this extends its access-history coverage. SRC-1 packages/coding-agent/src/session/ttsr-coordinator.ts:435-454; packages/coding-agent/src/session/session-context.ts:280-293; packages/coding-agent/src/sdk.ts:1678-1679.
Claims
| ID | Source-native claim and evidence | Conclusion status and boundary |
|---|---|---|
| CLM-1 | Coding agent with IDE integration and better tool outcomes; SRC-2 README.md:5-27,111-127 |
claimed. Runtime/tool wiring is supported; comparative percentages and universal model success are not verified by this run. |
| CLM-2 | First-class subagents, isolated worktrees, schema-validated yield and no sibling conflicts/orphaned edits; SRC-2 README.md:163-171 |
claimed. RTE-3/RTE-4 establish configurable branches, not an unconditional isolated/correct result. |
| CLM-3 | Advisor catches rushed mistakes and course-corrects; SRC-2 README.md:173-179 |
claimed. RTE-5/RTE-6 wire critique/delivery; detection accuracy and behavioral correction uninspected. |
| CLM-4 | Memory remembers codebase between sessions through retain/learn/recall and mental-model loading; SRC-2 README.md:213-215 |
claimed. Integrated memory routes establish scoped wiring, not reliable recall or benefit. |
| CLM-5 | Stream rules interrupt and survive compaction; SRC-2 README.md:155-157 |
claimed. RTE-10 and integrated replay routes support conditional delivery, not demonstrated correction. |
| CLM-6 | Stale Hashline edits are rejected and edit performance improves; SRC-2 README.md:205-207 |
claimed. RTE-11 includes recovery and warning branches, so literal universal rejection is too broad; outcome metrics remain unverified. |
| CLM-7 | Autoresearch keeps improvements, preserves correctness and avoids benchmark gaming; SRC-2 packages/coding-agent/src/autoresearch/prompt.md:20-45,99-103 |
claimed policy. RTE-8 measures, RTE-9 acts on supplied judgment; warnings and scope records do not enforce this policy. |
Evidenced absences
None recorded. Scoped unknowns and implementation branch limitations are represented as limitations, not source-wide negative searches. In particular, uninspected provider weights, remote memory internals and dynamic activation are not ABS records.
Behavioral-authority paths
| ID | Consumer | Channel | Force | Horizon and evidence |
|---|---|---|---|---|
| BAP-1 | Primary/worker model | System prompt and context in RTE-1 | Instruction/advisory content; no enforced truth | Current turn and retained replay. SRC-1 packages/agent/src/agent-loop.ts:1593-1639; packages/coding-agent/src/session/agent-session.ts:6303-6350. |
| BAP-2 | Registered tool executor | RTE-2 wrapper policy / UI | Enforcing allow/deny/prompt on that call | Invocation; explicit headless policies inherited. SRC-1 packages/coding-agent/src/extensibility/extensions/wrapper.ts:178-335; packages/coding-agent/src/task/executor.ts:944-985. |
| BAP-3 | Parent workspace merge | RTE-4 status/apply condition | Permissive admission of worker delta | Completed assignment; not semantic acceptance. SRC-1 packages/coding-agent/src/task/structured-subagent.ts:629-681. |
| BAP-4 | Primary scheduler/model | RTE-6 steering/aside/card | Enforcing scheduling interruption where selected; advisory note content | Live turn/next eligible turn or preserved card. SRC-1 packages/coding-agent/src/session/session-advisors.ts:1185-1288. |
| BAP-5 | Stream/model | RTE-10 abort plus hidden custom instruction | Enforcing abort; instructional retry content | Matched generation and retained injection. SRC-1 packages/coding-agent/src/session/ttsr-coordinator.ts:389-465. |
| BAP-6 | Experiment workspace and future experiment model | RTE-9 Git action and next-iteration status context | Enforcing commit/revert protocol; advisory metrics and instructions | Current experiment iteration/retained session. SRC-1 packages/coding-agent/src/autoresearch/tools/log-experiment.ts:125-200,341-379; SRC-2 packages/coding-agent/src/autoresearch/prompt.md:53-96. |
| BAP-7 | Hashline executor | RTE-11 content/anchor check | Enforcing edit rejection or permissive recovery | One edit against live/stored content. SRC-1 crates/pi-edit/src/modes/hashline/patcher.rs:176-250. |
BAP-20 — Advisory-memory authority. Local guidance asks for current-repo corroboration; Hindsight/Mnemopi snippets are background knowledge and current user/tool evidence takes precedence. This is intended knowledge authority, without a runtime guarantee that the model obeys that qualification. SRC-2 docs/memory.md:24-28,150; docs/mnemosyne-memory-backend.md:21-29; SRC-1 packages/coding-agent/src/hindsight/mental-models.ts:246-250.
BAP-21 — Learned instruction authority. Sharpshooter introduces retained decisions as instructions to follow unless overridden; selected managed skills and approved plans guide procedure. Admission controls include bounded/path-safe managed writes and sharpshooter evidence-substring/type checks. These controls constrain entry shape/provenance; they do not verify the truth or necessity of the resulting advice. SRC-1 packages/coding-agent/src/sharpshooter/backend.ts:117-125; packages/coding-agent/src/sharpshooter/extract.ts:252-310; packages/coding-agent/src/autolearn/managed-skills.ts:112-174; packages/coding-agent/src/session/agent-session.ts:5666-5691.
BAP-22 — Access-metadata authority. Branch/entry identities, memory scope/expiry, importance and embeddings programmatically select or rank delivered memory. They route/rank the evidence supply rather than validate its content. SRC-1 packages/coding-agent/src/session/session-context.ts:207-235,424-477; packages/mnemopi/src/core/beam/recall.ts:540-573,732-756; packages/coding-agent/src/hindsight/mental-models.ts:229-243.
BAP-23 — Retained material as learning input. The local extractor consumes persisted session history; the consolidator consumes retained extraction outputs; the autolearn capture model consumes the source agent's accumulated messages; Mnemopi extraction/sleep consumes stored content or its selected projection. These inputs have learning authority at those artifact-producing consumers because they shape later retained summaries, facts or skills. This differs from knowledge authority when the coding agent consults the finished memory. It establishes neither weight training nor observed improvement. SRC-1 packages/coding-agent/src/memories/index.ts:724-769,565-570; packages/coding-agent/src/sdk.ts:1214-1255; packages/mnemopi/src/core/beam/store.ts:306-338; packages/mnemopi/src/core/beam/consolidate.ts:972-1034; SRC-2 packages/coding-agent/src/prompts/memories/consolidation.md:1-30.
For BAP-20, the consumers are the coding model and, on recall-before-compaction branches, the summarizer; the channel is generated background/Memory Guidance context, the force is advisory policy, and the horizon is the current prompt plus later rebuilt prompts while selected content remains available. BAP-21 reaches the coding model as populated sharpshooter prompt files, selected skill bodies or approved-plan reinjection; instruction force is a prompt policy and lasts until override, replacement, deletion or different selection. BAP-22 reaches symbolic history/retrieval selectors through IDs, indexes and metadata; selection/ranking is programmatic protocol force on each assembly, under the storage/service contract. BAP-23 reaches extraction, consolidation and capture consumers through their model inputs or mechanical consolidation input; learning force is the dependency of a durable artifact on that input, with no guaranteed semantic improvement. Model consumers can disregard delivered content; none of these paths has observed activation evidence.
Runtime account
An ordinary omp request is translated into SDK session options, creating the agent with working directory, selected model, prompt, tool registry, credentials resolver and provider session/cache identities. The session prepares memory and context, lets configured extensions alter the turn, performs pre-prompt compaction, and enters RTE-1. Model output either returns text or requests tool work. Tool argument validation, wrapper policy and downstream domain checks precede the effect. Tool results enter the next model call; user steering and follow-ups can continue the run. Text mode emits the final assistant content or error and disposes resources; JSON mode emits events. Identity is principally local session/task identity and provider credentials; it is not a demonstrated multi-tenant security boundary. SRC-1 packages/coding-agent/src/main.ts:1-5; packages/coding-agent/src/sdk.ts:3531-3594,3716-3745; packages/coding-agent/src/modes/print-mode.ts:154-231.
Material alternatives change the guarantee boundary. Direct SDK/agent-core callers can provide their own stream function or hooks. Cursor native bridges use a separately constructed edit tool only if edit was granted, while extension same-tool native delegation inherits the caller's approval. Eval can execute code and re-enter tools; a bash pattern is not a rule over every process creation path. Native provider tools have remote effect boundaries; MCP/extension behavior depends on code outside this inspection. ACP adds client permissions with a distinction between default and explicitly configured yolo. Headless print mode ignores the interactive plan-on-startup default; headless workers intentionally override the approval mode. These are route-specific controls, not a universal isolation envelope. SRC-1 packages/coding-agent/src/sdk.ts:2840-2846,2902-2948; packages/coding-agent/src/task/executor.ts:975-985; packages/coding-agent/src/modes/print-mode.ts:134-151; SRC-2 docs/approval-mode.md:72-85,134-162.
Four static forcing cases were selected:
- A tool requires approval in a headless worker. Inherited explicit prompt/deny still reaches the wrapper despite the worker's yolo mode; no UI rejects a required prompt. This establishes an implemented failure branch, not an observed failure (RTE-2/RTE-3).
- An isolated worker exits successfully with a delta. Apply admission uses exit/error/abort and apply settings, while structured-output validity is separately recorded and upstream schema failure can produce a nonzero exit. Worktree success does not grant correctness warrant; non-isolated and apply=false branches remain material (RTE-3/RTE-4).
- A Hashline file hash is stale. Recovery/head-tail warning branches precede mismatch rejection. The guard reduces particular editing failures without proving correctness or universally rejecting stale edits (RTE-11).
- Autoresearch reports keep despite scope or metric disagreement. Code records warnings and still applies the keep branch; the model's judgment drives successor selection. The benchmark and declared checks are the relevant evaluator, not a built-in answer oracle (RTE-8/RTE-9).
The runtime serves open coding requests and an optional bounded experiment loop. In open work, model and user diagnose/propose; tools produce evidence, the model decides continuation, and configured user/tool/extension gates may veto effects. The advisor offers a separate model judgment. An expected answer is not supplied by the loop contract; task-specific tests or a user reference could supply one, but their actual authority and coverage are uninspected. In autoresearch, the goal/metric/benchmark are set up for that experiment; the model proposes code and logs keep/discard, code performs measurement and Git consequences, and the human can stop/reconfigure the run. The prompt tells the model to run correctness checks and preserve the benchmark across a segment; RTE-9 does not enforce correct measurement or honest judgment. No supplied expected-answer corpus was inspected. No autonomy grade follows from these roles.
Material revision admission is covered by RTE-2 (ordinary files/instructions/config), RTE-4 (worker changes), RTE-7 (capability loading), RTE-9 (experimental successors), RTE-10 (conditional rule injection), and the memory specialist's write/maintenance routes. Production releases and self-updater/evaluation machinery are uninspected. Parametric components are endpoint-selected; changing retained context or procedures can still implement learning without any proven weight update.
Execution-preflight disposition: no dynamic check planned. Considered a real coding task, headless approval probe, isolated worker merge and experiment keep/discard probe. Static branch inspection directly answers the admission questions. Dynamic checks would require a frozen runnable dependency/native build, isolated writable fixture, provider credentials/configuration for model routes and explicit execution scope; none was configured or invoked. No failed fetch or local validation command is treated as a target-system probe. Consequently this run establishes no actual tool safety, task success, faithful recall, latency or causal benefit.
Lens scoping
Memory/context scope
Full lens triggered by SRC-2 README.md:155-157,213-215, session retention and memory integrations. The fresh specialist owns source-native retained objects, write/read-back/maintenance routes, parameter roles and comparison profile. Scope and records are integrated below; static instructions and unchanged current-turn buffers are distinguished from accumulated memory. Opaque alternatives constrain complete-set classifications.
Epistemic scope
Full lens: SRC-1 RTE-1 through RTE-12 produce claims, execute checks or admit changes; SRC-2 CLM-1 through CLM-7 promise outcomes or policy. Assessed objects OBJ-1 through OBJ-10 plus adopted memory objects. Question: what does each check, retained artifact and acceptance-like transition actually warrant, and where can it change behavior? Deep provider/remote-memory internals, actual benchmark datasets and specialist product modes outside the declared boundary are unassessed. Whole-system scope requires these limits rather than a single epistemic grade.
Lens outputs
Memory/context lens
There are several memory systems inside this runtime. Session history preserves work continuity through branch selection and multiple kinds of compaction. Configurable backends add reusable experience. Autolearn can turn a completed tool-using turn into later lessons or procedural skills. Builtin autoresearch carries its own durable playbook and experiment results between iterations. These routes have distinct consumers and authority, so a single description such as “stores and recalls memories” hides material differences.
The backend selector has five values: off plus local, Hindsight, Mnemopi and sharpshooter. Sharpshooter is an additional active branch beyond the commission's README shorthand (SRC-1 packages/coding-agent/src/memory-backend/resolve.ts:6-26; SRC-2 docs/memory.md:3-11). Local memory turns older persisted sessions into a project summary, long document and optional skills. Mnemopi separates raw transcript content, user-only extraction text and marker-free embedding text. Hindsight sends retention work to an external service, then imports query results and long-running mental models. Sharpshooter keeps three project decision files, selected by a model using correction/regression/subtlety criteria.
Authority differs. Local memory, Mnemopi recall and Hindsight mental models are introduced as advisory context, with current user/repository evidence taking precedence. Sharpshooter instead instructs the agent to follow remembered project decisions unless the user overrides. Generated skills are instructions when selected. A wrapper's developer/system channel alone is therefore insufficient to assign the content's intended authority (BAP-20, BAP-21, BAP-22).
Context volume is managed at acquisition, maintenance and consumption. Local extraction caps each history input and limits scanned sessions; summary and lessons share an approximate startup budget. Mnemopi bounds query length, number of recalled rows and final prompt injection. Mental-model rendering divides a character budget and can drop trailing models. Compaction changes the history representation itself. In remote compaction the visible summary is a lead-in: the operative older context is opaque replacement history, not that summary (OBJ-24).
Write and maintenance findings
RTE-23 is the clearest offline source-session learning chain: persisted history → bounded extraction → stored per-session outputs → consolidation → future project guidance. “Raw memory” is already a model extraction; it is not the original trace. Phase-one schema/parse failures are rejected, common secrets are redacted, and empty output is treated separately. The phase-two document/skills writer can remove stale files. Its JSON schema and redaction control shape and some leakage; no cited semantic judge validates the generated claims (SRC-1 packages/coding-agent/src/memories/index.ts:724-807,957-1002; SRC-2 docs/memory.md:70-84).
RTE-25 has a materially different source partition. It stores a multi-author transcript but passes user-only text to fact extraction and a marker-free projection to embeddings/full-text access. This reduces assistant-authored content entering the extractor while still leaving assistant content in retained episodic evidence. Mnemopi sleep is not an LLM synthesis pass in the inspected implementation: it joins/encodes bounded content, sets llm_used: false, promotes it to episodic rows and retains source IDs. This supports consolidate/promote, not an automatic conclusion of new-claim synthesis (SRC-1 packages/coding-agent/src/mnemopi/state.ts:522-551; packages/mnemopi/src/core/beam/consolidate.ts:89-124,972-1048).
Maintenance over Mnemopi retained content includes duplicate detection/update of salience, working-row TTL/count forgetting, expiry/supersession without erasure, and later ranked exclusion. Explicit update/forget/invalidate operate on editable working/episodic rows; fact rows are read-only at the coding-agent interface. /memory enqueue drains extraction and requests consolidation; ordinary disposal can retain without fresh extraction or full promotion. No guarantee follows that the final turns finish all derived writes before the process exits (SRC-1 packages/mnemopi/src/core/beam/store.ts:233-264,436-461,644-663; SRC-2 docs/mnemosyne-memory-backend.md:35-40,174-193).
RTE-26 sends raw retained documents across an API boundary. Acceptance of an asynchronous request is not proof that extraction, consolidation or a mental-model refresh has materialized. Mental-model seeding is create-only; modifying seed defaults does not update existing server objects. The operator can request content refresh or delete/reseed, while /memory clear only clears local Hindsight state, not the remote bank (SRC-1 packages/coding-agent/src/hindsight/state.ts:359-371,459-480; packages/coding-agent/src/hindsight/mental-models.ts:35-39,142-175; SRC-2 docs/memory.md:152).
RTE-28 has an instructive admission split. Extraction validates that evidence is a literal substring of the current user prompt and that categories/friction fields have valid types. It does not require a friction flag to be true. The actual regression/subtlety/repetition law, nonduplication against project docs, newest-wins conflict handling and normative rewriting are model instructions in the consolidation prompt. They are wired judgments, not deterministic semantic gates. The successful writer consumes the delta files; the final instructions deliberately avoid source IDs and task chronology (SRC-1 packages/coding-agent/src/sharpshooter/extract.ts:252-310; packages/coding-agent/src/sharpshooter/consolidate.ts:177-194; SRC-2 packages/coding-agent/src/prompts/memories/sharpshooter-consolidate-system.md:5-43).
RTE-21 and RTE-22 also qualify as trace learning under the commissioned comparison contract: automatic history/report transformation, retained output, and later consuming context. Their purpose is continuation, so learning does not require novel knowledge. RTE-29 can produce global learned procedures; RTE-31 can revise an experiment playbook from iteration feedback. These have a stronger derivation route than mere raw trace storage. Approved plans and fired-rule logs remain ordinary authored guidance/access bookkeeping unless an additional extraction route is established.
Read-back findings
The named consumer and selector distinguish direction. In RTE-20, the runtime selects a branch and compaction boundary for the next coding-model call; that is push even if an operator selected the branch. In RTE-30, the model requests a specific memory or skill and receives it; that is pull. A file/row ID on that request supplies no identifier-based push claim.
Local startup supply is project-root identity selection followed by coarse availability and budget clipping. Summary is allocated budget first and lessons use the remainder; a large summary can crowd lessons out. Full documents and skills stay available through requested reads, so absence from injected context does not mean absence from storage (SRC-1 packages/coding-agent/src/memories/index.ts:182-226,1280-1281; SRC-2 docs/memory.md:30-40).
Mnemopi first-turn queries are built from the latest user prompt plus recent user-bounded turns. The wrapper asks each scoped bank, merges by IDs/content, sorts and limits. Shipped retrieval includes lexical and vector similarity with importance/recency weighting; optional graph/fact/temporal alternatives are documented. Defaults documented at the pin are 8 recall results, 3 context turns, 4,000 query characters and roughly 5,000 injection tokens. The compactor receives separately retrieved memory selected from its recent messages; this establishes delivery to the summarizer, not proof the summarizer preserves it (SRC-1 packages/coding-agent/src/mnemopi/state.ts:385-431,472-493; packages/coding-agent/src/mnemopi/backend.ts:138-145,316-318; packages/mnemopi/src/core/beam/recall.ts:732-756; SRC-2 docs/mnemosyne-memory-backend.md:44-62).
Hindsight has two push selectors. Query recall sends prompt/history, bank/tags/types and token/budget hints to an unknown external ranking implementation. Mental-model supply lists bank content, performs literal tag inclusion, sorts by name and applies a character budget; its default render budget is 16,000 characters. The latter is coarse and identifier-targeted selection, not embedding or semantic judgment inferred merely from the phrase “mental model” (SRC-1 packages/coding-agent/src/hindsight/state.ts:293-307,425-491; packages/coding-agent/src/hindsight/mental-models.ts:190-243,275-307).
Sharpshooter supplies populated project decision files on prompt construction, with a global injection budget and no question-dependent retrieval. Managed skills initially supply a discoverable catalog; the body remains a separate selection/read surface. Autoresearch automatically selects the active branch/session and a small result subset but injects the notes blob; no notes budget was established in the inspected hook. Approved-plan delivery similarly inlines the loaded body and retains its recovery path. No observed benefit, guaranteed activation or universal injection completeness follows from these wired paths.
Route completion, visibility and revision admission
The following rows extend the named canonical routes. Each row states the immediate result, later consumer, selection/expiry, effect and failure/recovery boundary. Execution effects are wired; semantic activation and successful recovery in a deployment are uninspected. Memory is scoped to the parent runtime/model consumers named here. A fresh worker's inherited memory, service bank and settings depend on its construction; propagation to every delegated worker was not traced, so no universal shared-memory visibility is asserted. Ordinary worker terminal results follow RTE-3, not an assumed shared memory bus. Static schema, size, path and substring checks are code protocols within their named writers; model curation criteria and instructions remain policies. Filesystem/SQLite/service durability remains an external contract.
| Route | Immediate result and later read-back | Selection, invalidation and admitted revision | Owner, rejection, recovery and terminal boundary |
|---|---|---|---|
| RTE-20 | Selected history returned internally to context construction; later coding-model invocation consumes it | Leaf/parent IDs, reset and latest compaction boundaries choose retained parts; branch changes select a different path | SessionManager owns reconstruction. No semantic revision is admitted by replay. Recovery uses retained entries/archive references; remote payload interpretation belongs to provider. Terminal result is assembled history. |
| RTE-21 | A compaction entry/continuation artifact; RTE-20 later supplies summaries, frames or opaque replacement history | Trigger/method and history/token boundaries choose source; previous summaries may feed new summaries; pruning discards selected detail | Model summarizer, mechanical archive/pruner or remote provider produces the replacement. Persisted-entry protocol admits it; semantic faithfulness gate unestablished. Original session entries and shake references afford recovery where retained; encrypted payload contents cannot be locally recovered as language. No blanket rollback guarantee inspected. |
| RTE-22 | Branch summary or rewind report retained at transition; RTE-20 exposes it after navigation | Navigation/checkpoint identity selects source and new position | Investigator or summarizer proposes language; runtime persists it. Branch navigation is operator-controlled; no semantic judge established. Recoverable context depends on retained old branch/checkpoint. Terminal result is transitioned branch with findings. |
| RTE-23 | Stored extraction then memory documents/assets; future startup gets summary/lessons and requested reads get full files | Project root, session eligibility, budgets and job leases; maintenance can replace/remove stale generated files | Phase-one model and phase-two model propose; parse/schema/redaction reject or constrain data. Generated writer admits content, not truth. Job state controls retries/leases; all-failure atomic recovery of multiple files uninspected. Startup refresh can change summary while keeping earlier lesson snapshot. |
| RTE-24 | Save confirmation/file change; next session snapshot supplies lessons | Normalization, redaction, duplicate bullet removal, newest-first maximum 100; non-bullet manual text preserved | Model invokes save or human edits file. Local bounds reject malformed/oversized fields; duplicate/count policy alters retained bullets. Recovery of dropped lessons requires external backups/history, not established here. |
| RTE-25 | Working-row retention/enqueue and optional extracted/episodic rows; first-turn/pre-compaction push or explicit pulls return bounded recall | Source cursor, bank scope, user-only extraction, marker-free embedding; TTL/count, salience, expiry/supersession and ranked exclusion | Wrapper/extractor writes; mechanical sleep promotes with source IDs. Update/forget/invalidate can revise working/episodic rows; fact rows read-only at coding interface. Full-row read before update is prompt policy. Explicit enqueue drains extraction and requests consolidation; disposal need not finish all derived writes. Internal transaction/race guarantees not established. |
| RTE-26 | Async retain request acknowledgement or requested recall/reflect response; hooks later supply facts to model/compactor | Bank/tags, full-session/sliding window and query budget; remote selector opaque | Host selects/submits; service decides internal extraction/selection. Acknowledgement does not establish materialization. Local clear preserves remote bank. Server rejection, retries and recovery beyond adapter contract uninspected. |
| RTE-27 | Listed/cached mental-model text; rebuilt coding prompt gets rendered snapshot | Bank, tag inclusion including untagged entries, name order, character budget, TTL | Seed is create-only; operator refresh/delete/reseed can replace remote objects, while reload only refreshes local snapshot. Service derivation/rollback opaque. Terminal local result may remain stale until refresh. |
| RTE-28 | Queued evidence-bearing delta, then replacement project instruction files; prompt rebuild supplies files | Committed prompt/event and extraction skip rules; all pending deltas/prior files/project docs inform consolidation; configured injection budget | Model proposes and judges friction/conflict; code checks literal substring and types, not that any friction flag is true. Successful writer consumes deltas; final files can omit source IDs. User override/clear and future rewrite can supersede instructions. Crash-atomic multi-file rollback uninspected. |
| RTE-29 | Private capture turn and save/manage-skill tool result; future catalog and selected body affect coding model | Enabled plus autoContinue/tool-completion threshold; abort/plan/goal exclusions; global discovery; authored same-name skill wins | Active-model copy proposes lesson/skill create/update/delete. Writer checks body/description, bytes and linked paths; no semantic acceptance oracle. Delete/rewrite are explicit revision/recovery surfaces; prior body/version recovery unestablished. |
| RTE-30 | Requested file/row/recall/reflect result to requesting model | Request path/ID/query and backend support; no extra push inferred from a requested ID | Read itself does not revise memory. Missing/unsupported result cannot guarantee recovery; local files/full Mnemopi rows afford recovering details omitted from previews. Backend mutation admission belongs to RTE-24–RTE-28. |
| RTE-31 | Notes replacement/idea append and logged results; next iteration system prompt receives selected records | Branch/session, current segment, baseline/best/last three and pending/violations; notes blob supplied without an established hook budget | Model writes playbook; store admits update. Measurement/keep correctness remains RTE-8/RTE-9. Reload reconstructs state across compaction/restart; overwritten note history recovery uninspected. |
| RTE-10 | Retained rule injection and fired names; later history rebuild/startup restores them | Branch history and TTSR match/repeat policy select already-used names | Coordinator persists/restores bookkeeping; it neither learns new rule definitions nor semantically approves them. Branch selection changes history. Abort/retry guard is the existing RTE-10 recovery path. |
| RTE-32 | Plan body/path added to next prompt; sent flag committed on delivery | Retained path identity, not in plan mode, unsent flag after reset/compaction | Runtime reloads selected file; prior plan approval supplies instruction authorization, not validated claims. Read failure/revision recovery beyond retained path uninspected. No extraction or learning claim from this reload alone. |
Trace-learning routes before aggregation
These mappings concern accumulated inputs, automatic durable transformations and later consumers, not model-weight training. No route has observed reuse success. Qualifying continuation transformations remain learning even where they are intended to preserve meaning.
| Route | Acquired source → durable derivative → later consumer | Reuse horizon | Timing relative to source work | Distilled form and uncertainty |
|---|---|---|---|---|
| RTE-21 | Session logs → summary/archive/remote continuation → coding model | Session-path continuation; task horizon not determinable | Online trigger, including operator-triggered automatic work; remote internal schedule unknown | Local language or mechanically image-encoded language with symbolic metadata; remote operative form opaque |
| RTE-22 | Session logs/tool investigation → branch summary/rewind findings → continuing coding model | Branch continuation; session ID does not settle task horizon | Online around branch transition/rewind | Natural-language findings; exact preservation/ampliation unobserved |
| RTE-23 | Prior session logs → stored extraction → project documents/skills → future project coding calls | Per-project | Offline relative to finished source sessions, staged extraction/consolidation | Natural-language summaries and mixed skill assets; new-claim synthesis unobserved |
| RTE-24 | Model-selected experience, potentially session/tool traces → lesson → future project startup | Per-project for local branch | Online explicit model save; future consumption | Language. Only trace-fed automatic saves qualify; literal human edits are authored memory, not this learning claim |
| RTE-25 | Session transcripts with distinct extraction/embedding projections → facts/episodes/access data → model/compactor | Configured project banks and cross-task global banks | Online retention, queued extraction and staged sleep; full terminal completion unobserved | Language plus symbolic facts/access structures; mechanical sleep is consolidation, not established ampliation |
| RTE-26/RTE-27 | Host session-log submission → remote memory/mental-model objects → model/compactor | Configured project/global bank reuse; independently populated corpus excluded | Async realization unknown | Returned strings inspectable; server derivation opaque. These routes do not establish trace learning independently of their opaque backend |
| RTE-28 | User-message event plus bounded history → queued decision → project instruction files → rebuilt prompts | Per-project | Online extraction plus staged consolidation | Natural-language instructions/provenance; factual entailment and new claims unobserved |
| RTE-29 | Accumulated source messages/tool traces → captured lessons/global skills → future model selection | Cross-task-capable global skills; lesson scope backend-dependent | Online follow-up capture, with later reuse | Natural-language procedures, potentially symbolic linked assets |
| RTE-31 | Experiment tool results/history → model-authored notes/playbook → next iteration | Research-session/goal/branch; mutable configuration prevents a universal task-horizon classification | Online iterative update | Language notes plus structured retained metrics; incorporation of any particular feedback unobserved |
Raw session retention, catalog discovery, fired-rule restoration and approved-plan reload alone do not qualify as trace learning. The complete known trace-source set classifies adapter-visible inputs to the established learning routes; it excludes undocumented server inputs. Unknown task horizons, remote forms and remote schedules prevent complete learning-scope, timing and distilled-form sets. Faithfulness testing remains not determinable because no retained dependence experiment entered this evidence boundary.
Comparison interpretation
The profile aggregates all included alternatives, not only the default off/local branch. The storage union uses the weakest afforded basis because Redis/SQL adapters document a SessionManager caller but no actual configured deployment was inspected. Service-object covers the remote interface only. Symbolic access metadata, natural-language content and opaque continuation payloads stay distinct; neither SQLite containers nor JSON wrappers establish the complete representational form of their contents.
Known lineage spans manual authored edits, foreign imports, mechanical archives/access structures and trace extraction. Stored embeddings are numerical access representations, not learned provider weights. No model-weight update is established inside the supplied boundary. Curation mappings are intentionally incomplete: reduction without new claims is consolidate; duplicate merging/detection over retained content is dedup; replacing a lesson/skill/decision is evolve; expiry while retaining history is invalidate; TTL/count/ranking loss is decay; working-to-episodic movement or raised importance is promote. reflect can synthesize an answer, but an answer without a retained later-consuming route does not establish retained-memory synthesis. Unobserved generated playbook content and external server behavior prevent a complete curation union.
Trace learning is known yes from several local chains without depending on Hindsight internals. Its remaining axes use the same routes. Source-session extraction and later project reuse establish per-project scope; globally discovered managed skills support cross-task reuse. Compaction, branch and rewind summaries prove later context use, but session identity does not fix task horizon. Autoresearch has an explicit experiment goal/branch, yet its configuration can evolve; this report does not equate every persisted research session with one immutable task. This prevents a complete learning-scope union.
The trace-source set covers the supplied host's captured inputs to these learning routes, including session-log submissions to Hindsight. It does not assert a complete source inventory for independently populated remote banks or undisclosed server-side learning. Those external acquisition/derivation routes are excluded, while their opaque returned objects remain included at the adapter boundary. Learning authority is supported separately by BAP-23: the same retained content may be knowledge for a coding-model read and learning input for a subsequent extraction/consolidation call.
Online continuation/capture, offline prior-session extraction and staged extraction/consolidation are supported partial timing mappings. External asynchronous realization remains unknown. Opaque remote continuation state prevents the distilled-form union just as it prevents the broader representational-form union. These unknowns are genuine classification limits rather than missing report sections, and do not block integration if preserved.
Epistemic lens
1. Source-and-claim boundary
See SRC-1/SRC-2 and the immutable boundary above. Assessed route families are runtime generation/tool admission, worker disposition, advisor feedback, TTSR, edit matching, retry, autoresearch and integrated memory transformations. Unassessed route families are the explicitly excluded services and incompletely traced specialist tools. CLM-1/2/3/4/5/6 are operational/benefit claims; CLM-7 supplies experiment policy. No general knowledge-production claim is inferred from the word memory or from successful tool execution. Missing run evidence prevents identifying any actual accepted candidate, behavioral activation or causal benefit. Canonical source anchors on each object and route apply to its sparse overlay; no overlay redefines its identity or evidence layer.
2. Epistemic-object inventory overlay
| Object | Truth-apt part and transformation question | Warrant limit |
|---|---|---|
| OBJ-1 | Explanations/proposed fixes can contain ampliative conjectures; individual outputs may instead restate acquired evidence | No instance observed; prompt completion does not establish truth. |
| OBJ-2 | Action request itself is non-truth-apt; text nested in arguments belongs to its destination object | Approval concerns permission/shape, not truth. |
| OBJ-3 | Acquired tool output; source/input lineage is the invoked tool | Warrant limited to that tool/domain and uninspected external source. |
| OBJ-4 | Code/config is executable content; embedded hypotheses or correctness claims are separate candidate propositions | Edit/merge checks do not validate semantic behavior. |
| OBJ-5 | Structured response may contain truth-apt findings | Schema and exit status cannot warrant those findings. |
| OBJ-6 | Advice may assert a defect or propose a better action | Separate model judgment, not expected-answer access. |
| OBJ-7 | Operational program/configuration, with no required truth-apt claim | Loading grants effects without semantic clearance. |
| OBJ-8 | Numeric observations acquired from benchmark output | Parsed value preserves the command's reported value; measurement construct validity uninspected. |
| OBJ-9 | Logged outcome/description can assert improvement; status is operational metadata | Caller-selected metric and status can disagree with captured measurement. |
| OBJ-10 | Normative instruction/trigger; rule body may include factual text but matching does not check it | Pattern occurrence warrants the match, not the rule's truth or suitability. |
| OBJ-20, OBJ-32, OBJ-36 | Acquired/imported messages or service-returned content; retention does not convert assertions into established facts | Source authors/provider warrant remain separate from faithful copying; no runtime instance observed. | | OBJ-22, OBJ-25, OBJ-26 | Generated continuation/navigation/rewind language; intended summary or report may preserve, omit or introduce claims | Transformation indeterminate without paired inputs/outputs; structural persistence is not entailment checking. | | OBJ-23 | Mechanical archive of linguistic source text plus symbolic shape metadata | Transformation is intended knowledge reshaping within selected bounds; omitted content and actual provider use unobserved. | | OBJ-24 | Opaque provider replacement history | Operative content and transformation not determinable; visible lead-in cannot substitute for the consumed payload. | | OBJ-27, OBJ-28, OBJ-29, OBJ-30 | Generated extraction, lessons and procedures; manual edits have authored lineage | Truth-apt assertions are indeterminate transformations; imperative procedures are non-truth-apt policy updates. Parse/redaction checks do not validate meaning. | | OBJ-33 | Mechanically bounded episodic content with provenance/veracity metadata | Reshaping and promotion, not an established new-claim synthesis or factual endorsement. | | OBJ-34, OBJ-37 | Extracted facts or evolving mental-model assertions | Generation/remote derivation indeterminate; source category, importance or model label is not warrant. | | OBJ-39 | Extracted decision supported by a literal user-prompt substring | Substring check establishes occurrence. It does not check entailment, future applicability or that friction criteria hold. | | OBJ-40, OBJ-41, OBJ-43, OBJ-46 | Instructions, procedures, playbook and plan, potentially containing factual claims | Direct instruction adaptation for imperatives; embedded truth-apt claims need separate candidate evidence. User authorization to act is distinct from semantic validation. | | OBJ-21, OBJ-31, OBJ-35, OBJ-38, OBJ-42, OBJ-44, OBJ-45 | Access, provenance, scheduling and ranking metadata | Non-truth-apt selection structures; a linguistic description may be a selection judgment, but use does not validate stored claims. |
3. Authority-route ledger overlay
All implemented rows below have architectural status implemented. All observed candidate states are no instance observed. Identity, endpoints, context, persistence and anchors remain owned by the cited canonical routes. Each functional row is distinct even where one runtime route performs several operations.
| Route | Function | Content/update relation | Check/transition target → evaluator/condition | Operational and epistemic force; limit |
|---|---|---|---|---|
| RTE-1 | content transformation | indeterminate; capable of ampliative conjecture | OBJ-1 → CMP-2 conditional generation | Candidate answer/tool plan enters use; no evidence-consuming acceptance. BAP-1. |
| RTE-1 | acquisition/return (other) | acquisition/import | OBJ-3 → selected tool output returned to model | May inform next action; warrant remains tool/domain-specific. BAP-1. |
| RTE-2 | check/evidence production | no content change | OBJ-2 effective arguments → policy/extension/UI conditions | Produces admission decision for this call, not evidence of factual correctness. BAP-2. |
| RTE-2 | operational admission/selection/consumption | non-truth-apt policy/content update when target is code/config | OBJ-2/OBJ-4 → successful admission | Allows effect; no correctness acceptance; ambient capabilities remain. BAP-2. |
| RTE-3 | check/evidence production | no content change | OBJ-5 → schema/status machinery where configured | Structural/termination status; payload truth not checked by schema. BAP-3 governs separate merge. |
| RTE-4 | operational admission/selection/consumption | no content change | OBJ-4 → apply flag and process outcome | Parent changes admitted; no evidence-consuming semantic acceptance. BAP-3. |
| RTE-5 | content transformation | indeterminate; often proposed defect conjecture | OBJ-6 → advisor model | Generates critique against supplied context; correctness uninspected. BAP-4 delivery separate. |
| RTE-6 | operational admission/selection/consumption | no content change | OBJ-6 → emission/duplicate filter and channel conditions | Scheduling interruption/continuation and advisory context; acceptance label here means delivery admission, not epistemic acceptance. BAP-4. |
| RTE-7 | operational admission/selection/consumption | non-truth-apt policy/content update: loaded capabilities/hooks | OBJ-7 → callable factory/load completion | Runtime behavior changes; queue rollback does not prove arbitrary side-effect rollback. |
| RTE-8 | check/evidence production | acquisition/import of measurements | OBJ-8 → experiment harness, output parser and exit status | Evidence only for that benchmark invocation; no automatic expected-answer oracle. BAP-6 subsequent use. |
| RTE-9 | disposition/acceptance | no content change | OBJ-4/OBJ-9 → caller status, pending-run requirement | Implemented operational keep/discard. Semantic evidence-consuming acceptance against a required criterion: not determinable from actual instances (none). Metric policy is advisory; no forced improvement check. BAP-6. |
| RTE-9 | retention | acquisition/import of caller metric/status and captured measurement | OBJ-8/OBJ-9 → logger/storage | Preserves discrepant fields and warnings; retention does not establish acceptance. |
| RTE-9 | behavior/policy adaptation | non-truth-apt update: commit/reset and successor selection | OBJ-4 → status branch | Kept code becomes successor; flags affect future comparison. Generalized improvement remains uninspected. BAP-6. |
| RTE-10 | check/evidence production | no content change | OBJ-10 condition against output → regex/AST matcher | Establishes a scoped match, not a semantic defect. BAP-5. |
| RTE-10 | behavior/policy adaptation | non-truth-apt instruction update | Matched rule → coordinator retry/injection | Aborts/discards/retries selected stream, model compliance uninspected. BAP-5. |
| RTE-11 | check/evidence production | no content change | OBJ-4 edit anchoring → hash/snapshot/recovery logic | Checks content correspondence within algorithm, not software truth or intent. BAP-7. |
| RTE-11 | operational admission/selection/consumption | non-truth-apt update: patch text | Candidate edit → admitted matching/recovery branch | Writes can proceed with recovery/warning; semantic validation separate. BAP-7. |
| RTE-12 | lineage/freshness/recovery | no content change, or removal of failed transient message | Failed response → classification and replay-state check | Retry/stop choice; does not prove exactly-once remote effects or answer quality. |
The memory overlay below has architectural status implemented at the host routes and observed candidate state no instance observed throughout. Opaque service-internal transformation/acceptance has architectural status not determinable; that separate limitation does not erase the implemented adapter.
| Route | Functional overlay and content relation | Target → evaluator/condition → transition | Operational force and epistemic limit |
|---|---|---|---|
| RTE-20 | acquisition/return; acquisition/import | Stored selected entries → identity/boundary selector → context | BAP-20/BAP-22 supply context; no evidence-consuming acceptance. |
| RTE-21 | content transformation; intended reshaping for mechanical archive, indeterminate for model/remote derivatives | History → summarizer/archive/provider → replacement entry | Retention plus later consumption. Faithfulness and candidate acceptance unestablished; remote internal architecture not determinable. |
| RTE-22 | content transformation; indeterminate summary/report relation | Branch history/investigation → model → branch findings | Retention and operational branch transition; an investigator report need not be an accepted explanation. |
| RTE-23 | content transformation; indeterminate claim relation | Session logs/extractions → extraction and consolidation models → project content | BAP-23 learning input; schema/parse/redaction checks govern admissibility, not semantic acceptance. |
| RTE-23 | retention and operational admission | Generated documents/assets → writer → replaced managed content | Future BAP-20/BAP-21 use; stale-file cleanup is maintenance, not refutation. |
| RTE-24 | behavior/policy adaptation or content transformation according to payload | Proposed lesson → normalize/redact/dedup/bound → retained lesson | Imperatives are non-truth-apt updates; factual generalization is indeterminate. Duplicate text checks concern existing bullets, not truth. |
| RTE-25 | acquisition and content transformation | Transcript/projections → storage/extractor → working/fact rows | Import then indeterminate extraction. User-only input narrows authorship, not truth. |
| RTE-25 | retention/lineage and operational promotion; intended reshaping | Eligible working content → mechanical sleep with source IDs → episodic record | Promotes retained content without an evidence-consuming correctness judgment; veracity aggregation does not supply a new oracle. |
| RTE-25 | lineage/freshness/recovery and selection | Retained rows → expiry/supersession/ranking/update constraints → selected recall | BAP-22 governs delivery; read-before-update remains instruction, not proven precondition. |
| RTE-26 | acquisition/return and retention request; acquisition/import | Host transcript/query → remote API → acknowledgement or returned facts | Implemented adapter, opaque internal derivation/admission. Acknowledged retention and requested answers do not establish materialized accepted knowledge. |
| RTE-27 | operational selection/consumption and maintenance | Listed remote text → local tags/name/budget/TTL → prompt snapshot | BAP-20/BAP-22; listing/reload is no content change. Server refresh transformation and semantic acceptance not determinable. |
| RTE-28 | content transformation; indeterminate claims, non-truth-apt decision directives | Prompt/context → extraction model → delta | Proposes learned decisions; no actual candidate observed. |
| RTE-28 | check/evidence production | Supporting string/typed fields → substring/type predicates → admitted delta | Evidence of occurrence/shape only; booleans may all be false. |
| RTE-28 | behavior/policy adaptation and retention | Prior files/deltas/docs → model judgment then writer → replacement instructions | BAP-21 changes future guidance. Model friction/conflict rules are policy; consuming deltas can reduce later provenance. No semantic acceptance established. |
| RTE-29 | content transformation and behavior/policy adaptation | Source messages → capture model/tool request → learned lesson/skill | BAP-23 inputs become BAP-21 instructions after selection. Path/size checks constrain artifacts; no correctness oracle. |
| RTE-30 | acquisition/return; acquisition/import | Model request → backend read/recall/reflect → response | Pull supplies content; generated reflect answer is not retained learning without a separate retained later-consumer route. |
| RTE-31 | content transformation and retention | Results/current playbook → model → revised notes | A proposed improvement explanation is indeterminate; procedural notes are direct adaptation. RTE-8/RTE-9 keep measurement and disposition separate. |
| RTE-10 | lineage/freshness/recovery; no new content induction | Fired-name history → rebuild/startup → restored manager state | BAP-22 routing history; replay does not discover or validate rules. |
| RTE-32 | operational consumption; no content change | Approved plan reference → mode/sent/path selector → prompt instructions | BAP-21 procedural force. Prior approval authorizes work; reload is not epistemic acceptance or trace distillation. |
RTE-13 is a classify-only direct policy adaptation: architectural status implemented, non-truth-apt effort selection, no content acceptance or retained learning claim. The consumer/channel is the primary model's effort option for this turn, with symbolic clamp force. Its model difficulty judgment does not supply an expected answer. See CMP-6 and RTE-13; no candidate lifecycle applies.
4. Per-object lifecycle disposition
OBJ-1, OBJ-5 and OBJ-6: transformation indeterminate for the unobserved payloads. Acquisition/restatement, derivation and ampliative conjecture remain possible. RTE-1/RTE-3/RTE-5 implement production; RTE-2/RTE-4/RTE-6 implement operational checks/admission/use. Lineage is prompt/context or delegated assignment, not guaranteed evidence support. No observed content permits a claim that preservation, entailment or ampliation occurred. Candidate-linked outputs, inputs and evaluative traces would resolve this. For any proposed defect/explanation that is ampliative, observation/conjecture/test/acceptance/integration all have observed candidate state no instance observed; generation architecture is implemented, actual evidence-consuming acceptance architecture for a supplied semantic criterion is not determinable within these generic loops. Operational delivery is not post-acceptance lifecycle integration.
OBJ-3 and OBJ-8: acquisition/import, discovery lifecycle not applicable to the direct recorded tool/measurement data. RTE-1/RTE-8 establish acquisition/lineage; warrant depends on source tool/harness. Any model explanation derived from these data is a separate candidate, not automatically warranted by an exit code.
OBJ-9: caller's improvement description is indeterminate; status itself is operational. RTE-9 preserves caller and measured values and changes code state. A claim of improvement would require a valid comparison, named criterion and candidate-linked keep decision. No instance observed at any phase. The architecture implements measurement (RTE-8), operational disposition and retention (RTE-9); prompt policy supplies an intended metric criterion, while semantic acceptance and post-acceptance integration remain not determinable without evidence of how the caller judged that particular experiment. A logged keep is not automatically a produced accepted knowledge claim.
No lifecycle record for OBJ-2: no candidate truth-apt output for this object; relevant update routes: RTE-2. No lifecycle record for OBJ-4: no required candidate truth-apt output in executable code itself; relevant direct-adaptation routes: RTE-2, RTE-4, RTE-9, RTE-11. Associated correctness propositions require their own candidate evidence. No lifecycle record for OBJ-7: no candidate truth-apt output required for this object; relevant update route: RTE-7. No lifecycle record for OBJ-10: no candidate truth-apt output required for the normative rule/trigger; relevant direct-adaptation route: RTE-10. Embedded factual assertions are not checked by matching.
OBJ-20, OBJ-32 and adapter-returned OBJ-36 are acquisition/import at the inspected retention boundary. Discovery lifecycle is not applicable to copying those inputs; truth claims within them retain their source's warrant. OBJ-23 and OBJ-33 follow mechanical reshaping routes (RTE-21/RTE-25): selected source text, archive/source IDs and bounded compilation establish lineage and intended preservation, without candidate-level proof of semantic fidelity. Discovery lifecycle is not applicable to that bounded mechanical operation. For OBJ-24, transformation and lifecycle disposition are indeterminate because the consumed provider payload is opaque; paired inspectable representations or provider evidence would resolve it.
OBJ-22, OBJ-25, OBJ-26, OBJ-27, OBJ-30, OBJ-34, OBJ-37 and the factual portions of OBJ-28/OBJ-29/OBJ-39/OBJ-40/OBJ-41/OBJ-43/OBJ-46 remain indeterminate transformations at the candidate level. Model summary/extraction/consolidation may reshape, derive or conjecture; exact input/output claims were not observed. For any ampliative claim, observation, conjecture, consequence derivation, test, acceptance and integration each have observed candidate state no instance observed. RTE-21–RTE-29/RTE-31 implement production, retention and consumption as specified, not a supplied end-to-end semantic acceptance lifecycle. Opaque Hindsight internal phases have architectural status not determinable. Occurrence checks, generated-file admission and later prompt delivery do not fill missing semantic test/acceptance/integration phases.
No lifecycle record for the non-truth-apt imperative portions of OBJ-28/OBJ-29/OBJ-40/OBJ-41/OBJ-43/OBJ-46: relevant direct policy/content-update routes are RTE-23/RTE-24/RTE-28/RTE-29/RTE-31/RTE-32. No lifecycle record for OBJ-21/OBJ-31/OBJ-35/OBJ-38/OBJ-42/OBJ-44/OBJ-45 as access/scheduling data: relevant selection/recovery routes are RTE-20/RTE-23/RTE-25/RTE-27/RTE-29/RTE-31/RTE-10. Their influence is operational selection rather than acceptance of a truth-apt candidate. These dispositions do not assert that every memory artifact lacks novelty; they state exactly what unobserved content prevents classifying.
5. System-claim versus route comparison
| Claim | Doctrine/design support | Implemented routes | Observed-run / causal support | Supported conclusion and mismatch |
|---|---|---|---|---|
| CLM-1 | README tool/performance claims | RTE-1/RTE-2/RTE-11 | uninspected / uninspected | Coding/tool substrate wired; comparative performance and universal success unverified. |
| CLM-2 | README subagent claim | RTE-3/RTE-4 | uninspected / uninspected | Optional isolation and structured completion machinery; no unconditional isolation, conflict freedom or semantic approval. |
| CLM-3 | README advisor claim | RTE-5/RTE-6 | uninspected / uninspected | Separate critique can interrupt/advise; better decisions not established. |
| CLM-4 | README memory claim | Adopted memory routes | uninspected / uninspected | Later delivery wired in scoped routes; faithful activation/improvement unverified. |
| CLM-5 | README stream-rule claim | RTE-10 and retained replay | uninspected / uninspected | Conditional interruption/persistence; unmatched semantics and compliance remain outside guarantee. |
| CLM-6 | README stale-edit/benefit claim | RTE-11 | uninspected / uninspected | Hash/recovery guard with warning branches; universal stale rejection overstates route. |
| CLM-7 | Experiment prompt policy | RTE-8/RTE-9 | uninspected / uninspected | Measurement and operational keep/discard; logger does not enforce metric honesty or improvement. |
6. Bounded conclusion
The runtime uses several distinct authorities: permission policies govern execution, hashes govern edit correspondence, schema checks govern output structure, advisor judgments govern interruption/advice, and caller experiment judgments govern keep/discard. None licenses the others' semantic claims. Tool evidence can support a task-specific conclusion when its source, interpretation and acceptance criterion are supplied; this static analysis observes no such completed candidate. Memory read-back and trace-fed learning describe retained influence, not truth or demonstrated benefit. No system-wide epistemic verdict follows.
Reconciliation
The fresh specialist used the frozen input digest recorded in Run identity and returned a complete report. Its proposals were registered as follows; these proposal labels are provenance mappings only, not a second operative inventory.
| Specialist proposals | Canonical disposition |
|---|---|
| MEM-CMP-1 through MEM-CMP-8 | CMP-20 through CMP-27 respectively; CMP-24 names the separate memory-author/capture role of the primary model, without asserting separate weights from CMP-2. |
| MEM-OBJ-1 | Split provisional content/access proposal into OBJ-20 and OBJ-21. |
| MEM-OBJ-2 / 3 / 4 | OBJ-22 / OBJ-23 / OBJ-24 respectively; linguistic summary, image-encoded archive and opaque remote payload remain distinct. |
| MEM-OBJ-5 | Split into OBJ-25 navigation summary and OBJ-26 rewind findings. Parent source check corrected their initially swapped line anchors: navigation uses packages/coding-agent/src/session/agent-session.ts:9569-9599 and rewind uses packages/coding-agent/src/session/agent-session.ts:8044-8091 (SRC-1); referents unchanged. |
| MEM-OBJ-6 / 7 | Split into OBJ-27/OBJ-28/OBJ-29 and OBJ-30/OBJ-31 respectively. |
| MEM-OBJ-8 / 9 | Split content into OBJ-32/OBJ-33/OBJ-34; access structures registered as OBJ-35. |
| MEM-OBJ-10 | Split returned documents, mental-model content and access into OBJ-36/OBJ-37/OBJ-38. |
| MEM-OBJ-11 / 12 / 13 | OBJ-39 / OBJ-40 / split OBJ-41 and OBJ-42. |
| MEM-OBJ-14 | New notes OBJ-43 and access OBJ-44, reusing existing experiment-output OBJ-8 and logged-status OBJ-9. |
| MEM-OBJ-15 | Reuse rule/injection OBJ-10; add distinct fired-history OBJ-45. |
| MEM-OBJ-16 | OBJ-46; no earlier canonical plan record needed merging. |
| MEM-RTE-1 through MEM-RTE-12 | RTE-20 through RTE-31 respectively. |
| MEM-RTE-13 / 14 | Amend existing RTE-10 for restoration; register RTE-32 for plan reload. |
| MEM-BAP-1 through MEM-BAP-4 | BAP-20 through BAP-23 respectively. |
All splits above were of local proposals before canonical allocation; no canonical identity was recycled. The RTE-10 amendment is attached to its existing record. Profile references were expanded to canonical parts while preserving the specialist's scope, values, evidence bases and explicit uncertainties. The local report is provenance; every substantive adopted mechanism and limit is retained in this result.
The specialist's integration issues are resolved: all eight component roles, sixteen object proposals, fourteen route proposals and four authority paths have destinations; sharpshooter is included among four active backends; opaque remote continuation and snapcompact remain distinct; autoresearch shares runtime experiment records; TTSR restoration extends the existing route; approved-plan reload has its own record; semantic gates are not inferred from shape/provenance checks; Hindsight clear versus server retention and uncertain selectors/forms/horizons remain explicit. The specialist's no-probe finding is a limitation, never an ABS record. Parameter invocations do not imply inspected weight identity or training.
After parent feedback the report qualified its local IDs, replaced an absence proposal with an uninspected limitation, added model component roles, explicitly included learning authority at retained-input consumers, and bounded trace-source classification to adapter-visible captured inputs. These revisions used the unchanged input and source pin. The final report digest identifies that corrected handback. No material integration issue remains unresolved. Shared source pointers and parent spot-checks are not claimed as independent convergent evidence.
Runtime and epistemic analysis share IDs; the local epistemic pass is not independent convergence. RTE-3/RTE-4 separate worker production from parent admission; RTE-5/RTE-6 separate critique from its operational force; RTE-8/RTE-9 separate measurement from selection. Source-native accepted advice, valid output and keep labels are not translated into epistemic acceptance. Documentation-versus-code narrowing is retained on CLM-2, CLM-6 and CLM-7. No status was upgraded to observed or causally supported.
Bounded synthesis
oh-my-pi is a coding runtime whose SDK connects model selection, context assembly, tool admission, effects and retained sessions. Its discriminating mechanisms are at the boundaries between those responsibilities: per-call approval rechecks revised arguments, optional worker isolation captures deltas, advisors can interrupt through selected channels, TTSR conditionally retries with instructions, and Hashline uses stored content to recover some stale edits. These are implemented mechanisms at a frozen source revision, not demonstrated outcome guarantees.
For ordinary coding work, model and tool breadth make configuration and host permissions consequential. A narrow advertised tool roster, a policy over bash and an isolated worker worktree are different controls. The default approval mode and headless worker override make parent authorization important; explicit per-tool deny/prompt still has force. A successful worker can make a delta eligible for integration while its substantive correctness remains a separate question.
For repeated work, session reconstruction and alternative compaction representations preserve continuity independently of the default-off memory backend. Local memory extracts older sessions into project guidance; Mnemopi separates transcript retention, fact extraction and recall access; Hindsight delegates retention/reasoning to a service; sharpshooter rewrites project instructions. Managed skills and research playbooks supply additional later procedural guidance. These distinctions determine what a future model receives and with what authority. Retention and distillation can change later context without any established model-weight update. For experimental work, the built-in benchmark/logging loop retains evidence and performs Git consequences, but trusts the model's keep/discard and metric report. Its warnings support inspection; they do not enforce truthful measurement or correctness.
A candidate-linked execution trace would change the operation assessment. A recalled-content intervention with controlled comparison would change the activation/benefit assessment. Inspectable remote-memory implementation would narrow profile uncertainty. An enforced semantic acceptance criterion, comprehensive process containment, or a changed merge condition would materially change the control assessment. Benchmark figures require retained comparable runs; present wiring alone cannot establish them.
Limitations
| Limitation | Affected records | Inspected boundary | Conclusion prevented | Resolving evidence |
|---|---|---|---|---|
| No runtime or causal evidence | SRC-1/SRC-2, all routes | Static committed code/docs | Observed correctness, activation, benefit, containment and performance | Frozen environment plus candidate-linked runs/interventions |
| Provider and external tool internals excluded | CMP-2/CMP-3, RTE-1/RTE-2/RTE-12 | SDK and bridge/adapters | Weight fixity, remote side-effect replay guarantees, tool truth | Exact provider/effect implementation and deployment contract |
| Deployed grants/settings absent | RTE-2/RTE-3/RTE-7 | Defaults and configurable paths | Actual permission/isolation envelope | Deployment config, extension set and OS boundary evidence |
| Selected code-route coverage | RTE-1 through RTE-12 | Included families stated above | Universal product guarantees or complete specialized-mode analysis | Focused inspection of omitted browser/desktop, LSP/DAP, prewalk, commit, security and host deployments |
| External performance evidence excluded | CLM-1/CLM-6 | README claims only | Comparative effect and causal attribution | Captured benchmark inputs, runs and intervention design |
| Current upstream revision unverified | SRC-1/SRC-2 | Locally present 2026-09-05 commit | Applicability to later changes | New source pin and new run |
| Remote memory and compaction opaque | CMP-26/CMP-27, OBJ-24/OBJ-36/OBJ-37/OBJ-38, RTE-21/RTE-26/RTE-27 | Shipped provider/service adapters only | Complete form, curation, selector, timing and faithfulness classifications; server materialization or deletion guarantees | Pinned remote implementation/contracts plus retained execution evidence | | No fixed task horizon for continuation | RTE-21/RTE-22/RTE-31 | Session, branch and mutable research-session consumers | Complete per-task/cross-task learning-scope union | Actual task boundaries and reuse traces | | Conditional/local memory coverage | RTE-23 through RTE-32, BAP-20 through BAP-23 | Named default/configurable routes, not every optional ranking branch or arbitrary extension | Actual deployment activation, universal delegated visibility, full injection and terminal write completion | Deployment settings, worker construction and end-to-end retained traces | | Model-authored memory content unobserved | OBJ-22/OBJ-25 through OBJ-30/OBJ-34/OBJ-37/OBJ-39 through OBJ-43, RTE-21 through RTE-31 | Source templates, schemas, selectors and writers | Claim preservation/entailment, semantic acceptance, complete synthesis classification and benefit | Paired candidate inputs/outputs, semantic checks and later-consumer intervention | | No inspected dependence test | All memory routes | Static source inspection; no runtime probe commissioned | Whether recall changed behavior or improved an outcome; no global absence of upstream tests asserted | Retained controlled recalled-content dependence experiment |
Verification and blockers
Semantic verification
Verified ordinary progression, alternate execution paths, four forcing cases, component endpoint/fixity distinctions, material revision admission, human/model/program roles, experiment mode and answer-oracle limits against the cited boundaries. Epistemic overlay separates architectural status from observed candidate state, checking from disposition and retention from integration. No static source is used as operation evidence. Canonical reference and source-anchor checks accompany final validation.
Memory integration checks covered RTE-20 through RTE-32 plus RTE-10 restoration: scope includes acquired/derived retained objects and their access structures, excludes static instructions and inaccessible server internals, and still includes opaque adapter-visible replacement history. Split object identities preserve those parts. Each automatic selector names its later consumer and selection input; requested rows/files remain pull. Each qualifying trace-fed derivative, including continuation and explicit lesson saves where trace-fed, has source/horizon/timing/form disposition before aggregation. The profile retains unknown complete sets for opaque forms, server selection/curation/timing, variable task horizon and unobserved dependence testing. Learning authority is attached to retained-input consumers separately from final advisory/instruction consumption. All fourteen axes preserve the specialist's assessments and use canonical record IDs.
The report run/source/pin/completion status and frozen input digest were checked. Parent source spot-checks confirmed backend defaults, sharpshooter admission and instruction wrapper, remote payload replay, autolearn conditions, local budget ordering, Mnemopi source partitions/mechanical promotion, Hindsight tag selection and approved-plan delivery. The split navigation/rewind anchor correction is documented in Reconciliation. Component roles distinguish callable model/service identity from unknown weights and upstream parameter changes. Structural validation and anchor range checks do not replace semantic review.
Deterministic validation
Validation target: commonplace-validate --full kb/reports/state/agentic-system-analysis/AAS-2026-09-05-oh-my-pi-02/result.md. PASS (clean): frontmatter, local links, memory comparison assessments/canonical references and type schema passed, with no warnings or failures. This is structural validation, not a target-system execution test.
Blockers
None.