Preliminary analysis: mining 45 connect reports (last two weeks) + kb/log.md

Method: three Sonnet subagents each read 15 full connect reports (batches split by directory: notes/agentic-systems/agent-memory-systems/reference; sources 1-15; sources 16-28 + work) and wrote structured raw findings to batch-1/2/3-raw-findings.md in this workshop. This file merges those three passes with the current kb/log.md, dedupes across batches, and gives a triage call per pattern. Nothing has been written back to kb/log.md or any library artifact yet — that's the next decision, not something this pass made unilaterally.

The dominant pattern, stated once

Every batch independently found the same shape: many unrelated sources converge as evidence on the same hub note, but the claim that explains why they converge is never the thing anyone writes down. cp-skill-connect runs one artifact at a time, so it's structurally the wrong place to notice this — it can only propose "this connects to that note," never "these six things you've now separately connected all imply an unnamed seventh claim." That's exactly the gap kb/log.md's SYNTHESIS category exists to catch, and this mining pass is evidence the backlog is bigger than the current log (24 lines, oldest entries from early June) reflects.

Three concrete instances, one per batch:

  • Batch 1 (library notes): runtime-structure-determines-governance-control-surfaces.md and agent-memory-is-a-crosscutting-concern-not-a-separable-niche.md independently deny "fourth component" status for two different whole-system functions (governance, memory) over the same scheduler/context-engine/execution decomposition — implying an unstated general claim that crosscutting functions don't add peer components.
  • Batch 2 (sources): six sources from disjoint literatures (credence-goods economics, agent-adaptation surveys, scientific-discovery benchmarks, off-manifold generalization theory, supply-chain security, technical-review comprehension) all land as evidence on the-boundary-of-automation-is-the-boundary-of-verification.md. Strong corroboration, but also a risk the note becomes a dumping ground rather than a sharpened claim.
  • Batch 3 (sources): SkillOpt, SkillRL, and the faithful-self-evolvers critique form an actual resolution chain — condensed/optimized experience risks being behaviorally inert unless the update loop couples the artifact to a validation gate or policy-training signal (SkillRL's own report says this explicitly) — but neither treat-continual-learning-as-substrate-coevolution.md nor deploy-time-learning-is-the-missing-middle.md (each hit by 4/15 reports in this batch alone) states the closure condition as its own claim.

Merged patterns and triage

# Pattern Evidence (batch: count) Recommendation
1 Hub-note convergence without an authored synthesis claim (see above) B1, B2, B3 — the pervasive pattern Append the 3 SYNTHESIS candidates above to kb/log.md now; the SkillRL/SkillOpt closure condition and the governance/memory crosscut claim both read as close to note-ready if picked up later
2 Reverse/bidirectional edges lag behind inbound citations B1: 9/15 reports lead with this Log as an ABSTRACTION entry; possibly worth a periodic "reverse-edge sweep" process, but that's a design question for its own workshop, not a one-line fix — flag, don't build
3 Missing .ingest.md companions for load-bearing sources B2: 4/15 (27%); B3: 3/15 (20%); B1 also flags 2 missing kb/sources/ snapshots Concrete and actionable now — 7 named sources need cp-skill-ingest runs (list below)
4 Terminology collisions misroute keyword search B1: computational-model tag conflates PL/naming concepts with scheduling/orchestration; B2: "scheduler"/"decompose"/"orchestrator" collide build-systems vs. agent-orchestration senses; B3: "skill" collides harness-invoked procedure vs. RL-trained behavioral memory (4/15 reports, explicit accept/reject split on the same two notes) Log all three as one ABSTRACTION-family observation each; the "skill" collision is the sharpest and most reproducible — two notes get accepted for one sense and rejected for the other in the same batch
5 Collection/label-vocabulary friction in connect itself B1: unauthorized complements label in use; no notes→sources "contrasts" label; no reference-collection label for a doc that generalizes another; B3: agent-memory-systems/COLLECTION.md not loaded in a single-collection connect run, stranding cross-collection edges in Off-authorisation Candidates Log as FIX entries; the cross-collection-scoping issue in B3 is a workflow limitation worth a note in kb/reference/proposals/ if it recurs beyond this one instance
6 Tag-README / index lag B1: computational-model-README.md missing a just-written note in 5/15 reports; B3: fresh snapshot absent from kb/sources/dir-index.md at discovery time Log as ABSTRACTION/FIX; both are the generated-index-lags-capture pattern already named in prior log entries (see kb/log.md line on mark-semantics) — this is corroborating evidence, not a new mechanism
7 Discoverability gaps on individual notes B1: skill-discovery-re-fires-in-every-sub-agent-context.md has tags: [] Direct one-line fix, not really a log entry — just add tags
8 Snapshot capture fidelity B2: mojibake/zero-width chars in an X-article capture; arXiv HTML capture missing PDF-only formula/table fidelity Log as FIX; low urgency, note for future normalization pass
9 Named-but-unhomed concepts B1: "privilege separation" (1 note, no definition); B3: "intent engineering," "update-time compute," "text-layer extended-cognition lineage" Watch items, not log entries yet — each is a single occurrence; log only if a second instance appears (matches the KB's own "needs more instances before naming" convention)
10 Coverage gap: asymmetric-information economics B2: dulleck-kerschbamer is the first source anchoring credence-goods/lemons/moral-hazard economics despite the verification-boundary note already leaning on labor economics Log as GAP, informational only

Ready-to-append kb/log.md entries (deduped, 10 total)

- SYNTHESIS: [runtime-structure-determines-governance-control-surfaces.md, agent-memory-is-a-crosscutting-concern-not-a-separable-niche.md, agent-runtimes-decompose-into-scheduler-context-engine-and-execution.md]: two independently-written notes both deny "fourth component" status for a whole-system function (governance, memory) over the same scheduler/context-engine/execution decomposition — implies an unnamed general claim that functions over a runtime crosscut its structural decomposition rather than forming peer components
- SYNTHESIS: [the-boundary-of-automation-is-the-boundary-of-verification, dulleck-kerschbamer-doctors-mechanics-computer-specialists, giants-generative-insight-anticipation-scientific-literature, interpolation-extrapolation-hyperpolation, in-toto-farm-to-table-guarantees, adaptation-of-agentic-ai-survey, can-llms-perform-deep-technical-comprehension]: six unrelated sources (economics, agent-adaptation survey, scientific-discovery benchmark, off-manifold generalization theory, supply-chain security, technical review comprehension) independently converge as evidence on the same note — the strongest cross-domain convergence hub found in this mining pass, and a risk the note becomes an evidence dump rather than a sharpened claim
- SYNTHESIS: [skillrl-evolving-agents-recursive-skill-augmented-rl source, skillopt-executive-strategy-self-evolving-agent-skills source, large-language-model-agents-are-not-always-faithful-self-evolvers source, treat-continual-learning-as-substrate-coevolution.md, deploy-time-learning-is-the-missing-middle.md]: skill-evolution sources converge on an unnamed claim — condensed/optimized experience risks being behaviorally inert unless the update loop couples the artifact to a validation gate or policy-training signal (SkillRL's own report states this explicitly for its SFT+GRPO loop); neither hub note states the closure condition as its own claim
- ABSTRACTION: [connect reports for full-identity-keys-decouple-a-batch-protocol-from-its-packing-axis.md, orchestration-needs-privilege-quarantine-not-permission-scope.md, runtime-structure-determines-governance-control-surfaces.md, skill-discovery-re-fires-in-every-sub-agent-context.md, vocabulary-collisions-prevented-at-write-time-not-read-time.md] — 9 of 15 connect reports sampled in one batch lead with "notes cited inbound get no return edge"; hub notes accumulate inbound citations from newly written notes faster than authors add matching reverse edges — a structural lag in the connect-to-authoring loop, not a per-note defect
- ABSTRACTION: [skills-are-instructions-plus-routing-and-execution-policy.md, skills-derive-from-methodology.md, problem-first-skill-inverts-solution-jumps connect, skillopt/skillrl connect reports]: "skill" is overloaded between harness-invoked procedural packaging (Claude Code / problem-first sense) and RL/optimization-trained agent behavioral memory (SkillOpt/SkillRL sense); both target notes get accepted for one sense and explicitly rejected for the other in the same batch of connect reports, with reports independently reaching for the phrase "title collision"
- ABSTRACTION: [build-systems-a-la-carte connect, a-new-way-to-think-about-composing-skills connect]: "scheduler"/"decompose"/"orchestrator" collide between build-systems vocabulary and agent-orchestration vocabulary in the same way "skill" does above — a keyword-only search pass will misroute on any of these three terms
- FIX: [kb/notes/COLLECTION.md, kb/reference/COLLECTION.md]: three connect reports each independently hit a label-vocabulary gap — an unauthorised `complements` label already in use, no notes-to-sources label for "similar-but-different" contrasts, and no reference-collection label for a doc that generalizes/refines a theory note's claim (plus two existing edges already misusing notes-collection labels on reference pairs)
- FIX: [large-language-model-agents-are-not-always-faithful-self-evolvers, prov-overview, we-should-take-text-optimization-more-seriously-2064027464926716154, borretti-human-routers-of-machine-words, build-systems-a-la-carte, in-toto-farm-to-table-guarantees, a-new-way-to-think-about-composing-skills-skill-g-2047124337191444844, semantic-engine.md]: eight cited-but-unpromoted sources across two batches have no `.ingest.md` companion (or, for semantic-engine, no `kb/sources/` snapshot at all) despite being load-bearing evidence for current notes; connect on each can only emit reverse-edge candidates — run `cp-skill-ingest` on the seven external sources and snapshot semantic-engine
- ABSTRACTION: [large-language-model-agents-are-not-always-faithful-self-evolvers connect off-authorisation section]: connect runs are single-collection-scoped by default, so the strongest edges for a source get stranded in Off-authorisation Candidates whenever the best target lives outside the collection whose rules were loaded (here: agent-memory-systems reviews of Dynamic Cheatsheet and Agent Workflow Memory) — worth checking whether this recurs enough to justify a routing convention
- GAP: [dulleck-kerschbamer-doctors-mechanics-computer-specialists]: KB has no coverage of asymmetric-information/mechanism-design economics (credence goods, lemons, moral hazard) despite already citing labor economics as verification-boundary evidence; this source is the first anchor

Dropped from the raw batch drafts as too report-specific to earn a durable log line (still true, just not log-worthy): the causal-inference-cluster saturation note, the computational-model-README.md lag (folded into pattern 6's index-lag observation instead of its own entry), individual snapshot-fidelity and citation-drift items, and single-occurrence vocabulary gaps (intent engineering, update-time compute, privilege separation) — these stay in the batch raw-findings files as a record but don't clear the "changes how someone builds or operates a KB" bar on their own.

Open questions for the user

  1. Append the 10 entries above to kb/log.md and delete the three batch-N-raw-findings.md scratch files (closing this workshop), or hold off?
  2. The 7-source .ingest.md backlog (pattern 3) is directly actionable — worth queuing cp-skill-ingest runs now, separately from the log-entry question?
  3. Three items (patterns 1's SkillRL/SkillOpt closure claim, the governance/memory crosscut claim, and pattern 1's automation/verification hub) read as close to note-ready rather than just log-worthy — worth a follow-up workshop, or should they sit in the log until someone picks them up organically?