Ingest report
Type: kb/types/type-spec.md
Authoring Instructions
Use ingest-report for source ingestion analysis. An ingest report is the
tracked, durable record of one URL-backed primary source and the KB analysis
derived from a local reading copy. It is not a copy of the source.
The primary reading copy lives under ignored kb/sources/.snapshots/. Its
path is local state and never appears in tracked frontmatter or links.
Assess fit relative to the installed KB's goals, local collection contracts, and current connection context. Interpret "our theory", "our stack", "our codebase", and "our practices" through those local goals and contracts.
Metadata
- Keep the H1 title text, including the
Ingest:prefix, to at most 100 characters. - Write
descriptionas a one-line retrieval filter between 50 and 250 characters. - Set
sourceto the canonical external URL of the primary source. - Set
capturedto the date or datetime of the observation used for the analysis, andcaptureto its capture mechanism. - Preserve
capture_scopefrom the snapshot when present. Do not upgrade a partial, abstract-only, or excerpt capture tofull-sourceduring ingest. - Set
genreto the primary source's evidential genre. - Set
snapshot_sha256to the lowercase SHA-256 of the exact bytes of the primary Markdown snapshot. The hash includes frontmatter, line endings, and the presence or absence of a final newline. It excludes companion files. - Never change
snapshot_sha256on an existing ingest. Changed source bytes are a distinct observation with a distinct snapshot basename and ingest path. - When the primary snapshot was mechanically derived from another retained
snapshot, set
original_snapshot_sha256to the lowercase SHA-256 of those exact precursor bytes. This gives the derivation input durable identity without treating a cache path as provenance or as a second primary source. - Use
type: kb/sources/types/ingest-report.mdfor the artifact type. - When the caller supplied an occasion — the question or job that brought the
source in, stated before it was read — record it verbatim in
occasion. Omit the field otherwise. The occasion governs only the selection sections named below; it is not evidence about the source. - Use
domainsfor two to four topic tags that make the report searchable. - Copy capture-adapter metadata such as
status_id,conversation_id,post_count, andapi_urlunder its existing flat field name. Do not copy a snapshot'stype,description,genre, ortags; set ingestgenrefrom the closer reading. - For a code-grounded ingest, add one
secondary_sourcesitem for every inspected implementation repository. Each item hasrole: implementationand a GitHub commit URL containing the full 40-character SHA. Do not record machine-local checkout paths. - Do not use the removed
original_snapshot,source_snapshot, orcode_revisionspath fields. - Link to durable KB artifacts and external sources in the report body. Never
link to the local
.snapshots/cache or generated connect reports.
Genre
The ingest's genre is the durable source classification. Use the vocabulary
and meanings in snapshot.md. A local snapshot may contain a
capture-time genre, but it is not authoritative. Set or correct the ingest
field after reading the source. The vocabulary is open: an off-list value warns
rather than fails. V1 has no operative local extension path that adds a known
value and its Limitations lens to this fixed ingest type. A recurring off-list
genre therefore remains warned until an ingest-side vocabulary mechanism is
adopted; a collection-local snapshot type does not extend this contract.
Sections
Classificationjustifies the source genre and identifies the author signal.Summaryis one paragraph for someone deciding whether to read the full source.Code Groundingis required whensecondary_sourcesis present. Link the reviewed revisions and pinned source files; distinguish mechanisms confirmed by inspection, experiment support artifacts that were present but not run, and claims that remain paper-only. State what code, if any, was executed.Quotesappears exactly once, immediately beforeConnections Found. It retains exact primary-source wording and human-resolvable locations so a later reviewer can judge ordinary source uses without the local snapshot.Connections Foundsummarizes the connection discovery findings and explains how the source fits the current KB, as compact prose naming the source's role (for example: anchor, technical basis, counterpoint, legal disposition, public statement, limitation) rather than a transcribed candidate list. Drop weak, speculative, or duplicate edges; keep only settled, durable judgments. If no casebook notes exist yet, say so plainly instead of substituting a full map of relationships to other already-captured sources, or framing the section as prospective connections for notes that do not exist yet. The generated connect report is working context only; do not cite it, link to it, or name its path in the ingest report.Extractable Valuelists three to seven items, ordered by reach and novelty relative to the installed KB's goals and existing KB connections. Whenoccasionis set, items bearing on it come first; if the source does not bear on it, one item says so.Limitations (our opinion)states where the source should not be trusted or over-generalized. Whencapture_scopeis notfull-source, state what the retained boundary prevents the ingest from establishing.Recommended Next Actionchooses one specific advisory next action. The ingest report recommends; it does not perform promotion. Whenoccasionis set, the action serves it unless the source does not bear on it.
Classification, Summary, Quotes, and Limitations (our opinion) are
observation sections. They follow the general contract regardless of
occasion, because later readers reuse them for jobs the occasion did not
anticipate.
Quotes Shape
Use this exact section when no source quotes have been retained:
## Quotes
No source quotes have been retained yet.
Use this shape for every populated item:
- **Source extract (verbatim):** <exact supporting content>
- **Source location:** <human-resolvable locator for that extract>
Use one or more adjacent Source extract (verbatim) / Source location pairs,
repeating both when support is non-contiguous. Copy exact snapshot text;
whitespace normalization lets one extract span wrapped lines. Do not put a
paraphrase, scope judgment, confidence assessment, limitation, or
target-specific transfer argument in this section.
Do not describe whether quotes are retained anywhere else in the ingest. That
state changes when the append-only Quotes pool grows. Use (snapshot required)
outside this section only for a specific claim that still needs broader context
than the retained extracts provide.
Extraction Standards
- Base extractable value on what is new relative to the connection context discovered by connect.
- Favor value that changes, supports, limits, or operationalizes the installed KB's current claims, decisions, policies, practices, or local domain work.
- Useful value classes include evidence for an existing claim, contradiction or limitation affecting current KB content, reusable method or workflow, data point or empirical result, vocabulary or framing that improves retrieval and discussion, operational warning or failure mode, and candidate artifact to write, update, retire, or review.
- Mark extractable value items with effort tags:
[quick-win],[experiment],[deep-dive], or[just-a-reference]. - Assess reach: high-reach findings explain why something works beyond the source's local context; context-bound observations should be flagged.
- Before writing limitations, ask what is surprising, what simpler account could explain the result, and whether the central claim is hard to vary.
- Be specific in the recommended action: name the note, reference document, runbook, instruction, policy, ADR, product requirement, dataset, incident note, or other local artifact to write, update, retire, or review. Filing as a source-only reference or scheduling a focused brainstorm are also valid when that is the right destination.
- Notes remain the default promotion target for transferable claims, but the recommended action may point to another local artifact type when collection contracts make that the better home.
Limitations Standards
Limitations (our opinion) is editorial judgment — label it as opinion. Name what is missing, cite a relevant KB note when one exists, and state what the gap means for the source's conclusions. The lens depends on the ingest's genre:
- Scientific papers — what was not tested: missing or naive baselines, limited benchmarks, configurations the literature or this KB already discusses, claims that do not generalize beyond the tested setup. Released source can confirm that a mechanism is implemented or expose its configuration, but static inspection does not reproduce training, benchmark, throughput, or quality results. State the remaining outcome-evidence gap.
- Practitioner reports — what is not visible: survivorship bias (what worked is reported, failed attempts are not), sample size of one, unacknowledged context such as team size, budget, or existing infrastructure.
- Conceptual essays and conversation threads — what is not argued: reasoning by analogy without testing whether the analogy holds, cherry-picked supporting examples, conflating naming something with explaining it, unfalsifiable framings.
- Tool announcements and design proposals — what is not shown: vendor bias and flattering benchmarks, missing failure modes or scaling limits, gaps between the announced design and real use.
- GitHub issues and code repositories — what is not durable: a single reporter's or author's view, point-in-time state that later commits may overturn, project history that records decisions without their later outcomes.
- Court opinions — what is not settled: interlocutory or preliminary rulings that later proceedings may overturn, jurisdiction-specific reasoning that may not generalize, procedural posture (for example, a motion to dismiss) that limits what the ruling actually decides.
- News articles and official statements — what is not independently verified: reliance on sources with their own interests, framing that reflects the outlet's or issuer's editorial stance, developing situations where later reporting may contradict early claims.
- Any other genre (the vocabulary is open) — fall back to the generic questions: what is surprising, what simpler account could explain it, whether the central claim is hard to vary, and what interests the author has in the framing. When a new genre recurs, add a dedicated lens here alongside its vocabulary entry in snapshot.md.
Template
---
description: "{one-line retrieval filter}"
source: {canonical external URL}
captured: "{date or datetime from snapshot frontmatter}"
capture: {capture mechanism from snapshot frontmatter}
capture_scope: {capture scope from snapshot frontmatter, when present}
genre: {source genre}
snapshot_sha256: {lowercase SHA-256 of the exact snapshot file bytes}
ingested: "{YYYY-MM-DD}"
occasion: "{caller's pre-reading question or job, verbatim; omit when none}"
type: kb/sources/types/ingest-report.md
domains: [{tag1}, {tag2}, {tag3}]
---
# Ingest: {source title}
## Classification
{Brief genre justification without repeating the field as a label.}
Author: {credibility signal}
## Summary
{One paragraph}
## Quotes
No source quotes have been retained yet.
## Connections Found
{Summary of connect discovery: which notes, what relationships, and what this source adds}
## Extractable Value
1. **{item}** -- {why it matters relative to existing KB connections}. [{effort}]
## Limitations (our opinion)
{Where this source should not be trusted or over-generalized}
## Recommended Next Action
{One specific action}
For a code-grounded paper ingest, add this frontmatter field:
secondary_sources:
- role: implementation
source: https://github.com/{owner}/{repo}/commit/{40-character-sha}
Add this section after Summary:
## Code Grounding
{Pinned repositories, claim-bearing source citations, inspection result, and execution status}