Next test: three pilots using generated citations

The operator commissioned this handoff on 2026-09-27 for execution in a new session. Its purpose is to test whether the citation-generation workflow in commit 5057c874 reduces quotation authoring errors in the same three pilots. This document prepares the test; no new analysis has been started.

Command for the next session

Paste this instruction into a fresh session at the Commonplace repository root:

Run the three-pilot citation-generation test described in
kb/work/agentic-memory-refresh/next-pilot-test.md. Act as the workshop
coordinator and delegate each analysis to a fresh source-only coordinator,
with its mandatory fresh memory specialist. Complete the three runs, audit
generator use and quotation failures, and record the comparison in the
workshop. Follow the fixed source pins and execution boundaries in that file.

Fixed inputs and baseline

Run Dynamic Cheatsheet, Mem0 and Napkin again, in that order, at these exact revisions. The fixed pins override the workshop's general upstream-refresh instruction for this test. Verify checkout origins and commit objects; read commit-addressed source without changing source worktrees.

System Repository Checkout relative to repository root Full commit Public destination
Dynamic Cheatsheet https://github.com/suzgunmirac/dynamic-cheatsheet related-systems/suzgunmirac--dynamic-cheatsheet 5cfe3c37e8e52b1d858d0f3df46e7f17c50991b9 kb/agentic-systems/reviews/dynamic-cheatsheet.md
Mem0 https://github.com/mem0ai/mem0 related-systems/mem0ai--mem0 94c3fe9f238f3dbf29c9ce98643bd71eb13077cd kb/agentic-systems/reviews/mem0.md
Napkin https://github.com/Michaelliv/napkin related-systems/Michaelliv--napkin 7582d6a46f5a11995956e60a59c41a5b242109f1 kb/agentic-systems/reviews/napkin.md

The workshop coordinator may read the previous trial and implementation acceptance. The previous trial had five failed source-check attempts, eight ambiguous quote blocks, one altered quote block, eight out-of-bounds citation occurrences, six bare-URL attribution errors and one copied-image-link error. Those categories overlap within attempts. All final publications passed; final success alone is not evidence of fewer authoring errors. The old verify-sources operation is gone, so compare defect categories and denominators as well as command failures.

The implementation passed 950 tests and the final publication cleanup passed 23 targeted tests. This next test evaluates real authoring, not another unit test run. At startup record HEAD, relevant code/instruction hashes, working-tree status and actual worker models. Check whether the producer differs from 5057c874; if it does, establish and document the intended test revision before launching workers. Do not silently test intervening producer changes.

Existing pilot reviews, retained results and workshop records include uncommitted work. Preserve it. Use guarded publication to replace only the three commissioned review destinations; allocate fresh, unused run IDs using the execution date. Do not reuse or overwrite any earlier run.

Execution and isolation

Follow the current analysis skill and its contracts. Run one coordinator plus its memory specialist at a time, reserving capacity for the specialist. Both lenses remain mandatory. Preserve the previous worker model/configuration where available and record differences; the previous six workers used gpt-6-astra.

Create each coordinator with fresh context (for the collaboration tool, fork_turns="none"). Supply only its source identity, checkout, pin, public destination, fresh run ownership and current method instructions. Tell it the purpose is a fresh source-grounded analysis under the current citation workflow. Explicitly supply repository doctrine if its runtime does not load it. Require a similarly fresh memory specialist under the current handoff contract. Do not give either worker this test document, earlier analyses, inventories, audits, findings, or failure examples. The workshop coordinator owns scheduling, recovery and the comparison; workers own their disjoint skill-defined outputs. Use completion events and owned output paths, not agent-status listings.

Authors must receive the current citation instructions through the skill and type contracts: use commonplace-quote, choose an occurrence, insert its citation unchanged, and assess semantic support themselves. One occurrence emits Markdown; two to ten emit candidates with metadata; more than ten asks for a longer selection. Ordinary navigation references omit ranges unless reusing generated locations. Do not add separate author quote checks or revive verify-sources. Publication uses the regular validator. Keep prepare/publish and the completed-run handoff checks prescribed by the skill.

Correctable draft failures may be repaired under the unchanged method; retain their evidence in worker traces. Do not fix producer code or instructions mid-trial. Record a discovered producer defect and its affected boundary; stop dependent work if it prevents valid completion. Never bypass validation. Prior analysis exposure or uncertain publication follows the skill's failure rule.

Evidence and acceptance

After each run, independently check its completed-run handoff. Audit all six worker traces, including nested tool results, exit statuses and stderr. Retain trace paths, hashes and model identities in a new dated cache directory under kb/reports/cache/agentic-memory-refresh/. Report inaccessible trace portions as audit gaps, not zero failures. Avoid creating a worker-side phase or retry ledger; the workshop audit summarizes the traces afterward.

For each coordinator and specialist, record:

  • Whether the citation instructions were loaded and commonplace-quote was used; successful calls, rejected requests, and reasons. Count requests above ten occurrences separately as expected requests for a longer selection.
  • Whether inserted citations match returned candidates unchanged, including citations copied from specialist to result. Record manual construction or alteration and cases where provenance cannot be established. This audits tool use; it is not another source-verification implementation.
  • Quotation and range errors reaching structural validation or publication: failed attempts, individual diagnostics, distinct affected passages and repeated diagnostics. Separate lookup/selection retries from malformed emitted citations, insertion mistakes, validator defects and schema failures.
  • Final quote counts by result, specialist report and compact review. Count the identical local and retained result only once. Report truncation and delegation problems separately, with recovery evidence where available.

Run the existing bounded matrix, table and statistics checks against exactly the three newly completed reviews, writing a fresh cache output directory. Inspect the existing scripts' help for their current interfaces. Verify their result identities and hashes. Update only the three inventory pointers after validated publication, preserving all other rows. A blocked pilot remains explicit; do not substitute an old result to make a three-row output.

Write a new dated citation-generation-rerun-<date>.md report in this workshop and link it from its README. Include run IDs, pins, producer start/end hashes, trace evidence, failure denominators, final validation, downstream checks, limitations and a comparison with the previous trial. Answer separately: did authors use the generator, did it emit valid citations, did authoring friction decrease, and did invalid quotations survive publication? Claim zero observed issues only within the inspected evidence; three stochastic analyses do not establish a general zero-error rate.

This commission covers these three analyses, guarded publication, bounded consumer checks, inventory updates and the workshop audit. It does not start the remaining corpus refresh, create a new landscape synthesis, change public comparison outputs, alter the output-document design, or authorize Git commits. Finish with the evidence-backed result or concrete blockers; do not expand the trial or repair the producer without a new instruction.