Ingest a paper with code
Type: kb/types/instruction.md
Use this conditional branch inside cp-skill-ingest. It replaces ordinary URL
snapshot resolution and adds code-grounding context to the normal connection,
drafting, validation, and reporting steps.
Keep the version-pinned paper snapshot as the primary source. Treat inspected code as corroborating evidence pinned to a Git commit, not as proof that the paper's experiments can be reproduced.
Resolve and capture the paper
- Extract the arXiv ID and any trailing version suffix such as
v2from the Papers with Code or arXiv target. - For both target kinds, fetch
https://paperswithcode.co/api/v1/papers/{arxiv_id}?include_resources=true. Use the response to discover paper metadata and associated repositories. Do not snapshot the Papers with Code page as the paper. - If that API is unavailable: for a Papers with Code target, extract the
canonical arXiv URL and
codeRepositoryvalues from the paper page's JSON-LD; for an arXiv target, take repository candidates from links in the paper or its abstract page. Treat either path as a resolver fallback, not paper evidence. - Resolve an unversioned target to the current arXiv version. Prefer the API's
versionwhen available. Otherwise inspect the arXiv PDF response'scontent-dispositionheader or the abstract page. Stop if no version can be established. - Set
paper_urltohttps://arxiv.org/abs/{arxiv_id}{version}and invokecp-skill-snapshot-webon it. Parse eitherSnapshot saved:orAlready snapshotted:to obtain the local primary snapshot. Retain its capture metadata and exact-file SHA-256 for the ingest.
Use https://arxiv.org/html/{arxiv_id}{version} as a reading and navigation
surface when it contains the full paper. The PDF-derived snapshot remains the
durable capture. If HTML is absent or incomplete, read the snapshot; consult
the versioned TeX source only to resolve extraction problems in equations or
tables. Do not substitute a Papers with Code summary or generated read endpoint
for the paper.
Resolve and pin the code
- Start with repositories returned by the Papers with Code API or the fallback resolver (JSON-LD or paper links).
- Prefer repositories marked official, but verify the association from primary evidence: the repository names the paper, arXiv ID, or released artifact, or the paper links back to the repository. Papers with Code is a discovery surface, not authority for the association.
- If several verified repositories implement distinct claim-bearing parts, keep each necessary repository. Do not select by star count.
- If the association is ambiguous, stop and ask which repository to trust. If no official or directly verified repository is available, report that a code-grounded ingest cannot be completed and offer the ordinary paper ingest.
Code grounding currently supports only GitHub repositories; if a verified
repository is hosted elsewhere, report that limitation and offer the ordinary
paper ingest. For each selected GitHub repository, normalize the URL to
https://github.com/{owner}/{repo} and use
related-systems/{owner}--{repo}/ as its checkout.
Before cloning, run git check-ignore -q related-systems from the main
project root; exit status 0 means the path is ignored. If it is not ignored,
stop and ask the user to approve an ignore rule rather than creating a large
untracked checkout. Create the directory when it is absent and ignored.
- If the checkout is absent, run
git clone "{repo_url}" "{checkout_dir}". - If it exists, verify that
originresolves to the same owner/repository. From inside the checkout, rungit fetch --all --prune, inspectgit status --short, and stop if the working tree is dirty. Otherwise fast-forward withgit merge --ff-only @{upstream}. - Stop rather than repurposing a checkout, overwriting local changes, resolving conflicts, resetting, or forcing an update.
Record reviewed_commit with git rev-parse HEAD and construct:
- revision:
{repo_url}/commit/{reviewed_commit} - file citation:
{repo_url}/blob/{reviewed_commit}/{path} - directory citation:
{repo_url}/tree/{reviewed_commit}/{path}
Inspect the implementation
Read the top-level listing, README, manifests, central implementation files, configuration, tests, training/evaluation scripts, and released result artifacts that bear on the paper's main claims. Classify each relevant claim:
- implemented -- source code realizes the claimed mechanism;
- artifact-supported -- configs, tests, scripts, or result files expose how a claim was operationalized without independently verifying its run;
- paper-only -- the checkout contains no evidence beyond restating the paper or README claim.
Inspect details that clarify or contradict the paper. Do not install dependencies, download weights or datasets, or run training or heavyweight evaluation during ingestion. Run a cheap existing test only when the environment is already ready and no download is required. Record exactly what, if anything, was executed.
Return to normal ingest
Continue at cp-skill-ingest Step 2 with:
- the version-pinned
paper_urlas the primarysourceand the local paper snapshot as reading input; - the snapshot's
captured,capture, flat adapter metadata, and exact-filesnapshot_sha256; - one
secondary_sourcesitem per pinned commit URL, each withrole: implementationandsource: {commit URL}; - the claim-to-code classifications and pinned file citations as drafting context;
- the execution status as both drafting and final-report context; and
- the paper version, checkout paths, and reviewed commits as final-report-only context.
Pass only the secondary_sources, claim-to-code classifications, pinned file
citations, evidence boundaries, and execution status to the worker executing
Draft an ingest report. Keep the paper version,
checkout paths, and reviewed commits in the parent as final-report context. The
worker's draft must include secondary_sources in frontmatter and add a ##
Code Grounding section after ## Summary. Link pinned revisions and source
files. State the claim classifications and execution status. Carry findings
into Connections Found, Extractable Value, and Limitations (our opinion)
where they change the judgment.
Do not cite related-systems/ paths or generated connect reports in the durable
ingest. Never describe source availability, static inspection, or passing unit
tests as reproduction of training, benchmark, throughput, or quality results.