Why the rerun still needed source-check repairs

The quote-matching change strengthened rejection, but did not finish aligning the authoring workflow with the new rule. Fresh memory specialists were not routed to the uniqueness requirement. Citation endpoints and some quote bytes were still reconstructed during drafting. A separate structural validator also rejected attribution syntax that the shared parser and authoring instructions accepted.

These are different causes. Increasing matching strictness cannot itself make the author supply the right occurrence, preserve every source character, or produce an attribution accepted by a second grammar.

Commission and evidence boundary

The operator requested this root-cause analysis after the three-pilot rerun. It explains the five failed verify-sources invocations under Commonplace commit 4a97ad715a4dbdcc6de09dea22417d1f88fbbbdf. Evidence is the recorded drafting commands, delivered tool outputs, diagnostics and corrections, plus the pinned source blobs and the implementation/instructions at that commit. No new system analysis, producer change or Git commit was performed.

The report distinguishes a directly observed construction error, an established contract mismatch, and a plausible behavioral contribution. It does not infer the model's private reasoning or claim a controlled intervention on its error rate. The existing audit retains the complete run IDs, source pins, final hashes and recovery history.

What failed

Failed invocation Stage reached Diagnostics Mechanisms below
Dynamic Cheatsheet specialist Source matching One quote occurring 32 times RC-1
Dynamic Cheatsheet coordinator Artifact validation, before source matching Six bare-URL attributions; one relative image link RC-3, RC-5
Mem0 specialist Source matching and ordinary source-anchor checks Seven ambiguous quotes; three out-of-bounds ranges RC-1, RC-2
Mem0 coordinator Ordinary source-anchor checks Two out-of-bounds ranges RC-2
Napkin specialist Source matching and ordinary source-anchor checks One changed quote; three out-of-bounds range diagnostics RC-2, RC-4

There are 23 citation/source diagnostics and one image-link diagnostic across five failed calls. They are not 23 independent mistakes: one formatting helper produced all six bare URLs, and Napkin reused the same bad end line across three anchors. Mem0's separate comparison-schema failure is outside this analysis. Availability probes and successful no-match searches are also excluded.

The distinction between stages matters. verify_sources first runs ordinary artifact validation, then quote matching and source-anchor checks. The Dynamic Cheatsheet coordinator was stopped before the matcher ran. Calling all five failures quote-matcher failures would misdiagnose that case.

RC-1 — The uniqueness rule did not reach the specialist's prescribed reading path

Established contract-delivery gap; likely contribution to author behavior.

Commit 6374109c added the unique-occurrence requirement to the main analysis skill and the quotation section of the result type. It did not update the standalone memory instruction or memory-report type. Those still explain that publication finds the text in the source, without stating that exactly one occurrence is required.

This omission matters because memory analysis deliberately runs in a fresh context. Its instruction explicitly directs the specialist to read only the main result type's Memory comparison fields and Status fields, alongside its own report type. The updated quotation rule is in a later section. The traces show that the specialists followed this restricted reading path:

  • Dynamic Cheatsheet read result-type lines 43–174 and 193–204.
  • Mem0 read lines 43–215.
  • Napkin selected the two named sections.

None of those reads includes the quotation rule at lines 247–266. The frozen inputs do not add a uniqueness instruction. All eight ambiguous quotes came from these specialists; neither coordinator introduced a new ambiguous quote that failed its result check. Some messages are encrypted in the traces, so this finding concerns the prescribed and recorded file-loading path, not proof that no off-band message could have mentioned uniqueness.

The source repetitions were real and explain why occurrence alone was insufficient. Dynamic Cheatsheet's short sentence appeared 32 times across nine JSONL rows: 24 times under steps and eight under final_cheatsheet. Mem0 duplicated five selected passages between synchronous and asynchronous implementations. Its SQLite initialization statement appeared in both classes' constructors and reset methods; another quotation appeared in both schema migration and creation.

The old presence-oriented instructions could be satisfied by these exact passages while the new uniqueness check rejected them. The contribution of the missing instruction is plausible, but these runs cannot establish how many failures updated guidance alone would prevent.

Required outcome: the specialist's actual required reading path must carry the same occurrence, normalization and range rules that its checker enforces. A coordinator knowing a rule does not supply it to a fresh worker. Prefer one explicitly loaded contract over independently maintained paraphrases.

RC-2 — Requested read bounds became unverified citation metadata

Directly observed for three bounds; generated inaccurate endpoints for the rest.

All eight range diagnostics were for ordinary source anchors, not quote blocks whose supplied range failed containment. Making quote ranges optional therefore did not remove this error surface.

The three Mem0 specialist endpoints reproduce its read-window endpoints:

Source Requested read ended at Actual final line Draft citation ended at
mem0/configs/prompts.py 1075 1062 1075
mem0/reranker/llm_reranker.py 190 173 190
mem0/utils/scoring.py 150 139 150

Those reads used sed -n without source line labels. Asking sed to print through a line beyond EOF succeeds and returns the available lines. Its exit status does not establish that the requested final line exists. The drafting code then wrote the same upper numbers into the report without deriving them from the actual blob length. This is an observable conversion of a selection request into an asserted source location.

Mem0's coordinator wrote 1069 and 774 for files ending at 1062 and 772. Its preceding reads requested other upper bounds, including 1070 and 775; the trace does not establish an exact copy rule for these two numbers. Napkin wrote 215 for a 198-line file in three anchors. Its initial read delivered the complete unnumbered crud.ts: the full blob is present inside the tool return even though that multi-file return carries a truncation warning. Missing source bytes therefore do not explain that particular error. The visible drafting commands emitted these endpoints as literal prose, without a source-derived endpoint calculation.

A contract conflict keeps this unnecessary metadata attractive: the main skill makes Git ranges optional navigation, while the result type's Status fields still asks for a full path and one or more ranges. The specialists did load that latter section. This is an established inconsistency, not proof that it caused each endpoint choice.

Required outcome: resolve whether ordinary ranges are optional. Omit unneeded ranges; when a range identifies evidence, derive and check its actual span against the pinned blob before emitting it. Do not treat a requested read window as the returned interval. Automatically clamping a bad endpoint would still require checking that the remaining passage supports the claim.

RC-3 — Structural validation retained a competing attribution grammar

Confirmed implementation/contract mismatch, not six missing sources.

Dynamic Cheatsheet's coordinator used a Python helper to copy source lines directly from pinned Git blobs. The helper ended each block with a bare full-commit GitHub URL. This agrees with the main skill and result type's stated URL alternative, and _attributed_citation parses the URL into a source, revision and range.

However, validate_quote_citations then independently applies _SOURCE_REF_RE, which recognizes only Markdown links and code spans. It reports that the bare URL names no source even though the shared parser has recognized one. The same helper produced all six rejected attributions. Wrapping them as Markdown links cleared these diagnostics without changing their repository, revision, range or quote text.

This is residual duplicate syntax recognition after introducing a shared parser. Telling the author to use Markdown links is a workaround. It does not resolve the disagreement between the documented input, parser and validator.

Required outcome: define the accepted attribution form once and make structural validation consistent with the parsed citation. If rendered links are required for a separate reason, say so explicitly and report a formatting error rather than claiming that a recognized source is absent.

RC-4 — A code comment was rewritten as prose inside a verbatim quote

Confirmed quote-body transformation.

Napkin's specialist copied two lines from bench/overview-exposure.ts but omitted their leading comment * characters. The words remained the same. The Git-source matcher intentionally normalizes whitespace only, so the modified passage did not occur in the blob. Restoring the two * characters made it pass.

This was not typography noise or an incorrect source revision. The quote went through a prose-style cleanup despite being represented as literal source text. The trace does not establish why the model removed the markers. Its instruction does require verbatim quotation, so this case cannot be explained solely by the missing uniqueness rule.

Required outcome: preserve the chosen source bytes when forming code quotations, including comment markers. Keep the author's interpretation outside the quote. Do not weaken code normalization to conceal transcription errors.

Confirmed interaction between exact copying and host-document validation.

Dynamic Cheatsheet's helper copied README lines 25–30, including an image whose target was figures/OverallPerformance.png. The copied text was faithful to the source. The report's Markdown link scanner includes links inside blockquotes, and its resolver interprets relative targets from the report's directory. It therefore looked for a local report-side image instead of the image in the pinned external repository.

This was an artifact link-health failure, not a source-text mismatch. Exact copying by itself cannot solve it: rewriting the target inside the quote would change the quoted source. The coordinator removed the unnecessary image line from the excerpt; the shortened passage still matched and supported the claim.

Required outcome: select only the necessary source passage. For a material relative link that must remain quoted, the validation contract must distinguish source content from links authored for the report. This does not justify ignoring ordinary broken report links or silently rewriting verbatim URLs.

What the evidence does not support

The current matching engine rejected the observed ambiguous or changed text as specified. No source change, wrong source pin, unavailable blob or publication race explains these five failures. Each failed source-check invocation was followed by a successful check after correction; the final artifacts remain verified. The procedure explicitly includes this correction loop.

Truncation and context volume can contribute to mistakes, but this trial did not isolate their effects. Napkin's complete crud.ts delivery rules out one simple truncation explanation. The restricted specialist reading path is independent of truncation: it deliberately skips the new quotation section. There is no basis here for attributing all failures to context size, model capability, or a need for another reviewer.

Five rejected calls versus two previously also does not show that authoring became worse. The checker now rejects ambiguities the earlier one accepted. Some new failures expose old contract inconsistencies. The supported outcome is stronger detection before publication, with remaining preventable repair work.

Repair priorities and a discriminating follow-up

  1. Align the actual contracts first. Route the complete quote contract to the fresh specialist; reconcile optional ordinary ranges; remove the disagreement over accepted URL attribution. These are demonstrated gaps.
  2. Preserve measured evidence during construction. Once the author chooses a passage, keep its exact text and actual source span together through rendering and integration. The citation helper already used in Dynamic Cheatsheet shows that exact copying is practical, but its accepted syntax also needs to agree with validation. The required outcome does not imply a new persistent evidence store or a larger workflow.
  3. Test the layer interactions before another full pilot. Exercise a valid pinned URL through both parsing and artifact validation, source-relative links inside retained quotations, a requested window beyond EOF, duplicate implementations, and literal code comments. Keep negative checks for a genuinely missing source, altered quote and ambiguous occurrence.
  4. Then measure fresh authoring separately from final acceptance. Hold the source pins and model setting fixed; count first-check failures by these categories and retain denominators. A clean final artifact measures the correction loop's success. Fewer first-check failures would support reduced authoring friction. Neither alone establishes semantic correctness.

These are proposed outcomes for follow-up work, not changes implemented or adopted by this analysis.

Trace map

Paths are under /home/zby/.codex/sessions/2026/09/27/. Numbers below are JSONL line numbers, not source-code lines. The rerun's trace index retains file hashes and model identities.

Trace Filename Relevant records
Dynamic Cheatsheet coordinator rollout-2026-09-27T15-36-59-01a0e315-4fce-7b71-8c14-09e9af772ceb.jsonl 215 copying/URL helper; 324 structural rejection; 331 correction; 334 pass
Dynamic Cheatsheet specialist rollout-2026-09-27T15-39-02-01a0e317-309d-70b3-9169-2f6f4a73fe66.jsonl 22/27 contract reads; 92 ambiguity rejection; 104 correction; 111 pass
Mem0 coordinator rollout-2026-09-27T15-54-34-01a0e325-68c7-72a2-b2ac-1b6efb2ba4cc.jsonl 115/162 source reads; 190 draft anchors; 353 rejection; 358 reread; 365 correction; 367 pass
Mem0 specialist rollout-2026-09-27T15-56-29-01a0e327-2a68-7c60-9d0e-c56967531d54.jsonl 22/30 contract reads; 75/84/94 read windows; 120 draft; 145 rejection; 160 correction; 161 pass
Napkin specialist rollout-2026-09-27T16-15-49-01a0e338-ddd6-7da2-ab3b-17f80cc42eef.jsonl 22 contract reads; 33/34 complete CRUD source delivery; 121 comment read; 135 draft; 151 rejection; 156 reread; 163 correction; 165 pass