Batch 1 raw findings — connect-report mining

Sources read (15 reports, full text):

  1. kb/agentic-systems/semantic-engine.md
  2. kb/agent-memory-systems/reviews/ReframeWeb.md
  3. kb/notes/a-derived-copy-of-recomputable-truth-must-be-checked-or-absent.md
  4. kb/notes/evidence/clausal-binding-scopes-a-captured-predicate-word-like-a-coinage.md
  5. kb/notes/formal-systems-can-assess-reach-through-causal-and-proof-obligations.md
  6. kb/notes/full-identity-keys-decouple-a-batch-protocol-from-its-packing-axis.md
  7. kb/notes/open-domain-memory-retention-needs-a-declared-output-spec.md
  8. kb/notes/orchestration-needs-privilege-quarantine-not-permission-scope.md
  9. kb/notes/reflection-may-improve-sample-efficiency-under-structured-shifts.md
  10. kb/notes/runtime-structure-determines-governance-control-surfaces.md
  11. kb/notes/skill-discovery-re-fires-in-every-sub-agent-context.md
  12. kb/notes/structure-inference-needs-capture-at-the-decision-surface.md
  13. kb/notes/theory-and-methodology-form-a-two-layer-execution-system.md
  14. kb/notes/vocabulary-collisions-prevented-at-write-time-not-read-time.md
  15. kb/reference/search-mechanisms-in-commonplace.md (report's own header links to a file titled where-change-candidates-come-from-in-commonplace.md — filename/frontmatter source: mismatch worth a quick check)

1. Synthesis Opportunities carried over

  • semantic-engine: Semantic Engine + CocoIndex + files-not-database + gap-plan imply a scoped operational DB/index layer belongs below semantic authority for high-volume ingest. Fairly specific — close to a design proposal, not yet written.
  • ReframeWeb: retrieval side-effects / read-purity as a design axis (ReframeWeb read-stamping + reasoning-bank embedding-append both break "reads are pure"). Only two instances — explicitly flagged as a watch item, still vague.
  • formal-systems-can-assess-reach: source + known-target-discovery + semantic-review + oracle-strength imply an evaluator taxonomy — formal-obligation verifiers vs retrospective target oracles vs semantic-judgment oracles answer different "does this generalize?" questions. Specific and well-formed, close to a note.
  • full-identity-keys: source + unified-calling-conventions imply "a well-chosen stable interface key converts what would be a protocol/architecture choice into a free policy choice made by the caller." Specific mechanism claim, close to a note; explicitly deferred ("do not author during connect").
  • open-domain-memory-retention: source + elicitation-requires-maintained-question-generation + evaluate-memory-by-effects imply "a declared need is the shared prerequisite for auditable coverage across the memory lifecycle" (recurs at admission, elicitation, evaluation). Well-formed, cross-cutting, close to a note.
  • runtime-structure-determines-governance: source (governance) + agent-memory-is-a-crosscutting-concern (memory) both deny "fourth component" status over the same scheduler/context-engine/execution decomposition, implying a general claim: functions over a runtime crosscut its structural decomposition rather than forming peer components. Well-formed, two clean instances, good note candidate.
  • orchestration-needs-privilege-quarantine: three artifacts (dynamic-workflows, GBrain, Secure LLM-Wiki) now independently instantiate the read/act privilege split; report concludes the source note itself already is the synthesis — no separate note needed.
  • All others (a-derived-copy, clausal-binding, reflection-sample-efficiency, skill-discovery, theory-and-methodology, vocabulary-collisions, structure-inference, search-mechanisms) report None — either the cluster is already an articulated argument chain or the note itself is already the synthesis.

2. Recurring cross-report themes

  • Missing reverse/bidirectional edges is the single most common finding type. 9 of 15 reports lead with "notes cited inbound get no return edge": a-derived-copy (4 inbound-only notes, one load-bearing), ReframeWeb (2 hand-curated evidence: lists that should cite it), open-domain-memory (kb-goals direction-of-primacy question), skill-discovery (skills-are-instructions bidirectional gap), vocabulary-collisions (2 bidirectional gaps), full-identity-keys, theory-and-methodology, runtime-structure, orchestration-privilege-quarantine. Suggests hub notes accumulate inbound citations faster than authors add matching return edges — a structural lag in the connect→author loop, not a per-note defect.
  • kb/notes/computational-model-README.md (selective, no complete: mark) is flagged as missing a just-written note in 5/15 reports: full-identity-keys, orchestration-privilege-quarantine, runtime-structure, skill-discovery (blocked — see below), vocabulary-collisions. The tag/README pair is a chronic index-lag point.
  • llm-context-is-composed-without-scoping recurs as the grounding hub for "scope must be imposed architecturally" across 3 unrelated claims: injection-defense (orchestration-privilege-quarantine), skill-discovery-leak (skill-discovery, already linked), and vocabulary-scope-as-schema-position (vocabulary-collisions). A genuine conceptual hub note.
  • codify-versus-llm-decision-heuristics pulled in as a new edge from two independently-written notes: formal-systems-can-assess-reach and theory-and-methodology-form-a-two-layer-execution-system. Second confirmed hub.
  • self-improving-systems-README.md's complete: true mark is used as a positive navigation shortcut (skip the by-tag rg sweep) in 3 reports (formal-systems, reflection-sample-efficiency, search-mechanisms) — working as designed, evidence the ADR 026 mark mechanism pays off in practice.
  • Missing kb/sources/ snapshots for material a note leans on: semantic-engine (no snapshot of the Semantic Engine repo itself) and reflection-sample-efficiency (8 arXiv papers cited as bare inline URLs, zero captured as ingests).
  • Collection label-vocabulary gaps, each independently discovered: a-derived-copy finds an existing edge using complements, not in kb/notes/COLLECTION.md's authorised set; orchestration-privilege-quarantine wants contrasts toward a kb/sources/ target but only evidence/derived-from/see-also are authorised there; search-mechanisms finds kb/reference/ has no label for a "this doc generalizes/refines that claim" relationship, and that two existing edges already misuse grounds/evidence (notes-collection labels) on reference→reference/reference→notes pairs.

3. Systemic/maintenance issues

  • Collection authorization gap, agentic-systems → instructions: kb/agentic-systems/COLLECTION.md does not authorize outbound links to kb/instructions/, even though semantic-engine's ingest analysis maps directly onto cp-skill-ingest/ingest-directory. (semantic-engine)
  • Collection authorization gap, notes → sources: no contrasts-equivalent label exists for "structurally similar external system, different driver" cases; only evidence/derived-from/see-also are authorised toward kb/sources/. (orchestration-privilege-quarantine)
  • Collection label gap, reference collection: no authorised label for a reference doc that generalizes/refines a theory note's claim; compounded by two already-committed edges misusing notes-collection labels (grounds, evidence) on reference-collection pairs. (search-mechanisms)
  • Unauthorised label in use: complements on an existing edge from a-derived-copy-of-recomputable-truth-must-be-checked-or-absent.md, outside kb/notes/COLLECTION.md's label set. (a-derived-copy)
  • Discoverability gap — empty tags: skill-discovery-re-fires-in-every-sub-agent-context.md carries tags: [], making it invisible to every tag-README and by-tag rg sweep despite being a structured-claim note with real content. (skill-discovery)
  • Tag overload: computational-model tag conflates two loosely related bodies — PL/naming concepts (scoping, homoiconicity, vocabulary collision) and scheduling/orchestration architecture — flagged as a candidate for a covered_by split. (vocabulary-collisions)
  • Named-but-unhomed concept: "privilege separation" appears in exactly one note KB-wide despite being a transferable, named security principle — no definition or tag-README head yet. (orchestration-privilege-quarantine)
  • Local vocabulary loans with zero reuse: physics terms (matching, cutoff, correspondence — effective field theory) and cognitive-architecture terms (proceduralization, ACT-R, SOAR) in theory-and-methodology-form-a-two-layer-execution-system have no other KB occurrences; flagged as fine for now but a drift risk if they recur without a definition.
  • Possible missing synthesis trait: theory-and-methodology-form-a-two-layer-execution-system carries 8 outbound edges and several independent assertions, reads like a synthesis-style note, but does not declare the trait. (theory-and-methodology)
  • Edge-density ceiling in practice: reflection-sample-efficiency's note already had 7 outbound edges and had passed a "connection-inflation" review gate; the connect run deliberately held back 3 more one-hop-reachable candidates to stay conservative — shows the inflation gate is actively shaping connect output, not just note-writing.
  • Classification-boundary flag: semantic-engine sits in kb/agentic-systems/ by user direction but isn't a whole agentic runtime; report suggests a subcategory or split collection may eventually be needed for "ingest infrastructure below the KB" analyses.
  • Repeated seedling/load-bearing-label caution: at least 3 reports (full-identity-keys, orchestration-privilege-quarantine ×2, skill-discovery) self-flag a proposed grounds (load-bearing) edge landing on a status: seedling target, recommending the author confirm the premise or downgrade to see-also/extends first.

4. Candidate kb/log.md entries

  • - SYNTHESIS: [runtime-structure-determines-governance-control-surfaces.md, agent-memory-is-a-crosscutting-concern-not-a-separable-niche.md, agent-runtimes-decompose-into-scheduler-context-engine-and-execution.md]: two independently-written notes both deny "fourth component" status for a whole-system function (governance, memory) over the same scheduler/context-engine/execution decomposition — implies an unnamed general claim that functions over a runtime crosscut its structural decomposition rather than forming peer components
  • - SYNTHESIS: [open-domain-memory-retention-needs-a-declared-output-spec.md, elicitation-requires-maintained-question-generation-systems.md, agent-memory-requirements/evaluate-memory-by-effects.md]: the same "criterion must come from a declared spec, not the incoming stream or stored existence" move recurs at three separate memory-lifecycle stages (admission, elicitation, evaluation) without a named connecting claim
  • - SYNTHESIS: [formal-systems-can-assess-reach-through-causal-and-proof-obligations.md, known-target-discovery-benchmarks-show-reachability-not-discovery.md, semantic-review-catches-content-errors-that-structural-validation.md, oracle-strength-spectrum.md]: three distinct "does this generalize?" evaluators (formal-obligation verifiers, retrospective target oracles, semantic-judgment oracles) are used across the KB without a taxonomy distinguishing them
  • - ABSTRACTION: [connect reports for full-identity-keys-decouple-a-batch-protocol-from-its-packing-axis.md, orchestration-needs-privilege-quarantine-not-permission-scope.md, runtime-structure-determines-governance-control-surfaces.md, skill-discovery-re-fires-in-every-sub-agent-context.md, vocabulary-collisions-prevented-at-write-time-not-read-time.md] — 5 of 15 connect reports sampled independently flag kb/notes/computational-model-README.md as missing a just-written note; the selective/non-complete tag-README pattern is lagging note production on this tag specifically
  • - FIX: [kb/notes/COLLECTION.md, kb/reference/COLLECTION.md]: three separate connect reports (a-derived-copy-of-recomputable-truth-must-be-checked-or-absent, orchestration-needs-privilege-quarantine-not-permission-scope, search-mechanisms-in-commonplace) each independently hit a label-vocabulary gap — an unauthorisedcomplementslabel in use, no notes→sources label for "similar-but-different" contrasts, and no reference-collection label for a doc that generalizes/refines a theory note's claim (plus two existing edges already misusing notes-collection labels on reference pairs)
  • - FIX: kb/notes/skill-discovery-re-fires-in-every-sub-agent-context.md: carries tags: [] and is therefore invisible to every tag-README and by-tag rg sweep despite being a normal structured-claim note; needs computational-model and/or architecture tags
  • - DUPLICATION: [semantic-engine.md source ingest, reflection-may-improve-sample-efficiency-under-structured-shifts.md]: two connect reports independently find load-bearing external material with no kb/sources/ snapshot (the Semantic Engine repo; 8 arXiv papers cited as bare URLs) — pattern of source-tracking lagging note authorship, not a one-off