Batch 2 raw findings — mining 15 kb/reports/connect/sources/*.connect.md

Sources mined (frontmatter source:, report date): adaptation-of-agentic-ai-survey-post-training-memory-skills.md (06-09) · a-new-way-to-think-about-composing-skills...skill-g-2047124337191444844.md (04-23) · borretti-human-routers-of-machine-words.md (06-14) · build-systems-a-la-carte.md (07-06) · can-llms-perform-deep-technical-comprehension.md (07-17) · causal-inference-using-invariant-prediction.md (07-16) · causal-learn-causal-discovery-in-python.md (07-16) · dowhy-expressing-and-validating-causal-assumptions.md (07-16) · dulleck-kerschbamer-doctors-mechanics-computer-specialists.md (07-06) · emergent-analogical-reasoning-transformers.md (05-26) · externalization-in-llm-agents-unified-review.md (04-13) · giants-generative-insight-anticipation-scientific-literature.md (04-24) · how-we-built-our-knowledge-base-2077822555159945507.md (07-17) · interpolation-extrapolation-hyperpolation.md (05-19) · in-toto-farm-to-table-guarantees.md (07-06).

1. Synthesis opportunities carried over

  • adaptation-of-agentic-ai: "adaptation evaluation must be component-counterfactual, dynamics-aware" — connects oracle-strength theory + reliability-dimensions + evaluate-memory-by-effects. VAGUE (not yet a claim, just a topic).
  • a-new-way-to-think-about-composing-skills (×3): (1) composition-depth degradation = flat-context dynamic-scope mechanism; (2) hierarchical composition and hierarchical context scoping are the same move at different layers; (3) skill-tier testing cost tracks the topology→isolation→verification chain's verification-cost step. All three WELL-FORMED (named notes, sketched mechanism/thesis) but explicitly deferred pending the source's still-missing .ingest.md.
  • borretti-human-routers-of-machine-words: "delegating prose generation to an LLM skips the concretization work that turns vague ideas into committed ones." WELL-FORMED, single clean claim, 4 named anchor notes, would host the Weizenbaum citation.
  • build-systems-a-la-carte: "the KB's derived-artifact freshness machinery is a build system" (scheduler = staleness-due logic, rebuilder ladder = dirty-bit→verifying-trace→content-address). WELL-FORMED, 4 concrete artifacts mapped to paper's classification; workshop kb/work/lineage-mechanisms/verification-locus-and-provenance-theory.md already drafting this.
  • can-llms-perform-deep-technical-comprehension: disagreement-preservation is an information-retention guarantee, not an adjudication guarantee. WELL-FORMED, sharpens an existing note's language.
  • causal-inference-using-invariant-prediction / causal-learn / dowhy: none beyond the already-created formal-systems-can-assess-reach-through-causal-and-proof-obligations.md. (All three say this — see §2.)
  • dulleck-kerschbamer: "agent-produced knowledge artifacts are credence-good-like" — maps verifiability/liability/breakdown onto representational-form, document-types-should-be-verifiable, the review-gate system. WELL-FORMED but flagged as needing to clear the "changes how someone builds a KB" bar before promotion.
  • emergent-analogical-reasoning-transformers (×2): (1) "analogy is structure transfer, not similarity search" — WELL-FORMED, 4 named artifacts; (2) "cognitive analogies need an operationalization ladder" — VAGUE, "may be useful for future reviews."
  • externalization-in-llm-agents (×2): (1) "externalization reframes context engineering as cognitive burden relocation" — WELL-FORMED, has a thesis sentence, 4 notes, but flagged as resting on several seedling notes; (2) "protocolization is externalized coordination guarantee" — WELL-FORMED, secondary.
  • giants: (1) "backcast benchmarks turn discovery into target reconstruction" — WELL-FORMED, names a reusable pattern; (2) "parent selection is the hidden hard part" — WELL-FORMED.
  • how-we-built-our-knowledge-base: enterprise-memory claim pairing this source with Databricks — MODERATE (says it's "close to" an existing ingest recommendation, undecided between a note vs. lightweight system review).
  • interpolation-extrapolation-hyperpolation: hyperpolation as the missing bridge between discovery/reach notes and oracle-boundary notes (generating an off-manifold candidate vs. evaluating it). WELL-FORMED, 5 named artifacts.
  • in-toto: three-source cluster ("verify a chain of production steps") with build-systems-a-la-carte + prov-overview. Explicitly speculative, "needs the ingest report first" — VAGUE/deferred.

2. Recurring cross-report themes

  • Hub: the-boundary-of-automation-is-the-boundary-of-verification.md. Hit by 6/15 reports (adaptation-of-agentic-ai, can-llms-perform-deep-technical-comprehension via technical-constraints, dulleck-kerschbamer, giants, interpolation-extrapolation-hyperpolation, in-toto), each from an unrelated domain (economics of credence goods, causal-boundary theory, agent-adaptation surveys, scientific-discovery benchmarks, off-manifold creativity, supply-chain security). This is the batch's strongest convergence point — external evidence keeps landing on the same "verification cost is the automation boundary" claim from disjoint literatures, which is a strong corroboration signal but also a risk of the note becoming an evidence dumping ground.
  • Hub: automated-synthesis-is-missing-good-oracles.md / oracle-strength-spectrum.md. Hit by giants and interpolation-extrapolation-hyperpolation directly, and indirectly by the causal cluster and can-llms via reach/technical-constraints notes. Same oracle-theory neighborhood as above, one layer more specific (synthesis/generation rather than automation generally).
  • Causal-inference cluster converges exactly. causal-inference-using-invariant-prediction, causal-learn-causal-discovery-in-python, and dowhy-expressing-and-validating-causal-assumptions — ingested same day (07-16) — all land on the identical two targets (reach-assessment definition, formal-systems-can-assess-reach-through-causal-and-proof-obligations.md), are pairwise compares-with each other, and each explicitly reports "no synthesis beyond the already-created note." Reads as a single ingestion campaign that has already saturated its target note; a fourth causal-discovery source would likely add little without broadening the note's claim.
  • Hub: "discovery is seeing the particular as an instance of the general." Hit by emergent-analogical-reasoning-transformers, giants, and interpolation-extrapolation-hyperpolation — three independent domains (neural analogy mechanisms, scientific-insight benchmarks, geometric generalization) converging on the same discovery-as-recognition note.
  • Missing .ingest.md blocks the outbound deliverable. borretti, build-systems-a-la-carte, in-toto, and a-new-way-to-think-about-composing-skills (4/15, 27%) are snapshots with no matching ingest report; each report can only produce reverse-edge candidates and has to punt its "Connections Found" to a hypothetical future ingest. This is a structural backlog, not four unrelated observations.
  • Keyword-collision rejections recur. build-systems-a-la-carte explicitly rejects the agent-orchestration "scheduler" cluster as vocabulary collision (build-scheduler vs. agent-scheduler); a-new-way-to-think-about-composing-skills does the analogous rejection for "decompose"/"orchestrator" against agent-runtimes-decompose-.... Suggests "scheduler"/"orchestrator"/"decompose" are overloaded terms across the KB's computational-model and agent-orchestration areas that a keyword-only rg pass will misroute.

3. Systemic/maintenance issues

  • Missing ingest reports (backlog): borretti-human-routers-of-machine-words, build-systems-a-la-carte, in-toto-farm-to-table-guarantees have no .ingest.md at all; a-new-way-to-think-about-composing-skills... likewise. All four reports say the fix is a single cp-skill-ingest run.
  • Snapshot capture fidelity, two distinct failure modes by source type: (a) how-we-built-our-knowledge-base — X/article capture picked up runs of zero-width/mojibake characters; visible content usable but flagged for a future normalization pass. (b) emergent-analogical-reasoning-transformers — arXiv HTML capture (not full PDF), so formula- and table-level details should be checked against the source before precise technical claims are promoted.
  • Stale/thin derived-note hygiene: research/adaptation-agentic-ai-analysis.md (flagged by the adaptation-of-agentic-ai report) cites a bare external arXiv URL instead of the local snapshot/ingest report, breaking local lineage tracking, and carries only type: in frontmatter — no description, status, traits, or tags — so it won't surface in description/tag scans.
  • Search-tooling reliability: externalization-in-llm-agents-unified-review reports two qmd semantic-search failures (GPU/rerank-context errors) during discovery; the report explicitly flags its search trace as not a complete semantic pass. Single occurrence in this batch but worth tracking if it recurs elsewhere.
  • Durable collection-coverage gap: dulleck-kerschbamer (economics of credence goods) notes the KB has zero coverage of asymmetric-information/mechanism-design economics (no Akerlof lemons, moral hazard, adverse selection) despite the-boundary-of-automation-is-the-boundary-of-verification.md already leaning on labor-economics evidence — this source is the first anchor in that area.
  • Uncorroborated concept flagged as a gap, not a candidate: in-toto notes its "graceful degradation under partial key compromise" (k-of-n functionary thresholds) has no matching KB note; flagged as a future link target if/when the KB develops a claim about redundant/independent verification degrading gracefully (parallel to error-correction-works-above-chance-oracles-with-decorrelated-checks).

4. Candidate kb/log.md entries

  • SYNTHESIS: [the-boundary-of-automation-is-the-boundary-of-verification, dulleck-kerschbamer-doctors-mechanics-computer-specialists, giants-generative-insight-anticipation-scientific-literature, interpolation-extrapolation-hyperpolation, in-toto-farm-to-table-guarantees, adaptation-of-agentic-ai-survey, can-llms-perform-deep-technical-comprehension]: six unrelated batch-2 sources (economics, agent-adaptation survey, scientific-discovery benchmark, off-manifold generalization theory, supply-chain security, technical review comprehension) independently converge as evidence on the same note, making it the strongest cross-domain convergence hub found in this mining pass.
  • SYNTHESIS: [causal-inference-using-invariant-prediction, causal-learn-causal-discovery-in-python, dowhy-expressing-and-validating-causal-assumptions, formal-systems-can-assess-reach-through-causal-and-proof-obligations]: all three same-day causal-inference source ingests converge on one note and report no further synthesis opportunity — the note has likely saturated this source cluster.
  • GAP: [dulleck-kerschbamer-doctors-mechanics-computer-specialists]: KB has no coverage of asymmetric-information/mechanism-design economics (credence goods, lemons, moral hazard) despite already citing labor economics as verification-boundary evidence; this source is the first anchor.
  • FIX: [borretti-human-routers-of-machine-words, build-systems-a-la-carte, in-toto-farm-to-table-guarantees, a-new-way-to-think-about-composing-skills-2047124337191444844]: 4 of 15 mined connect reports (27%) target snapshots with no matching .ingest.md, so their outbound Connections Found could not be authored — a standing ingest-report backlog surfaced by connect-report mining.
  • FIX: [adaptation-of-agentic-ai-survey-post-training-memory-skills]: research/adaptation-agentic-ai-analysis.md cites a bare external arXiv URL instead of the local snapshot and carries only type: in frontmatter (no description/status/traits/tags) — breaks local lineage tracking and won't surface in scans.
  • FIX: [how-we-built-our-knowledge-base-2077822555159945507, emergent-analogical-reasoning-transformers]: two distinct snapshot-capture fidelity gaps this batch — X/article captures carrying mojibake/zero-width characters, and arXiv HTML captures lacking formula/table fidelity versus the full PDF.