Check comparison and synthesis coverage
Historical initial pilot, before ADR 093 adopted per-value evidence. The results, contracts and counts below belong to that frozen first trial.
What must remain possible
The historical kb/agent-memory-systems/systems.csv contains 155 rows and 55
columns. The live reader uses 14 normalized axes with separate assessment,
evidence basis, and canonical-record references. These contracts differ;
column-name equality and unchanged counts are not acceptance criteria.
The following map comes from the historical CSV header and the current
src/commonplace/lib/systems_matrix.py contract, not historical system findings.
| Historical field family | Current evidence field | Acceptance query |
|---|---|---|
storage_substrate |
storage_substrate |
Decode JSON value sets; count membership once per eligible system. |
representational_form, form_* |
representational_form |
Natural language, symbolic, and parametric membership. |
lin_* |
lineage |
Authored, imported, trace-extracted, and other-compiled membership. |
auth_* |
behavioral_authority |
Separate knowledge, instruction, enforcement, routing, validation, ranking, and learning. |
wa_* |
write_agency |
Manual and automatic membership. |
op_* |
curation_operations |
Consolidate, dedup, evolve, synthesize, invalidate, decay, promote. |
trace_learning |
trace_learning |
Count qualified write routes; do not translate this field into improved capacity. |
ts_* |
trace_source |
Session logs, tool traces, event streams, trajectories. |
df_* |
distilled_form |
Natural language, symbolic, parametric derivatives. |
ls_* |
learning_scope |
Per-task, per-project, cross-task. |
lt_* |
learning_timing |
Online, offline, staged. |
read_back_direction, rb_pull, rb_push |
read_back_direction |
Pull and push membership and overlap. |
sig_* |
read_back_signal |
Coarse, identifier, inferred lexical/embedding/judgment selection. |
rb_faithfulness_tested |
faithfulness_tested |
Preserve uncertainty separately from supported yes/no. |
read_back_notes |
Axis rationale and canonical result routes | Read retained evidence for qualitative explanation; no standalone CSV notes column. |
public_repo, clone_path, one_line |
source_identity, source register, one_line |
Source identity and retained bytes replace a local checkout dependency. |
The producer also records new identity fields: public-review and exact-result hashes, run, revision, cutoff, boundary, tier, and memory comparison scope. Those are prerequisites for comparing refreshed results without blending different source versions or silently counting repeated runs twice.
Checks to execute after publication
- Pass the identical explicit review list to
scripts/build_systems_matrix.py,scripts/render_systems_table.py, andscripts/analyze_matrix.py. Write trial artifacts separately from the full-corpus comparisons. - Verify matrix values, assessments, bases, and canonical IDs against every selected result; verify the table records the same input identities.
- Recompute selected memberships and overlaps from a frozen synthesis bundle. Eligible positive counts require code-grounded, known values at wired, observed, or causally supported basis. Report exclusions beside denominators.
- Run summary generation through
synthesize-agent-memory-landscape. Preserve the bundle identity and executable query/output ledger. Qualitative claims require reading full retained records, not only CSV fields. - Test rejection of a missing or ineligible review and preservation of the previous output bytes on failure. Do not modify a published input for a test.
The numerical analyzer computes normalized entropy over complete value sets. That is not the same statistic as entropy over each old boolean column. A membership query can recover the old kind of binary count from an eligible value set, but the old and new analyzer outputs must not be presented as a longitudinal numerical series without reconciling these units and evidence filters. This pilot checks output capabilities, not historical number equality. Its mutual-information function returns zero for fewer than five paired rows. A three-system pilot therefore checks that path executes, but cannot establish the validity or usefulness of its redundancy diagnostics. Likewise, no document-only case is selected initially: tier separation is part of the contract, but needs a later document-grounded pilot.
Executed checks
On 2026-09-26, both matrix builder and table renderer were called with a missing
main review and, separately, a legacy review path. All four calls exited 1.
The errors distinguished file absence from not a main-review path. Each test
started with an existing temporary output containing previous-output; its
bytes were unchanged after the rejected call. No published input was modified.
The unbounded uv run python scripts/analyze_matrix.py also exited 1 on
2026-09-26: kb/agentic-systems/reviews/pond.md: missing or mismatched retained
result; regenerate the main review. This is an observed corpus prerequisite,
not a defect in the current pilot targets. The explicit pilot population must
not be presented as a full-corpus rebuild, and the current public comparison
files are not thereby certified fresh.
The first positive smoke check selected only the completed Napkin review. The
matrix builder, table renderer, and numerical analyzer all exited 0; the first
two emitted one code-grounded row to the ignored pilot cache. The analyzer
reported weaker afforded assessments separately from wired values. This
establishes one-result compatibility, not the final three-result trial or
equivalence to historical classifications.
Selecting the Napkin review twice was also rejected (exit 1, no output file):
multiple selected reviews of source https://github.com/Michaelliv/napkin;
choose one boundary explicitly. Repeated selection cannot inflate the system
count through that path.
All three pilot analyses have now published and passed checked handoffs.
The first synthesis bundle preparation failed because Napkin's retained result
links three ontology definitions not included by the default capture. Repeating
the command with the supported --ontology arguments for theory-builder.md,
reflective-system.md, and self-improving-system.md succeeded. This is a
required input-discovery step for future bundles, not grounds to strip valid
links from an immutable analysis. No script change was needed.
The explicit three-result bundle was frozen before interpretation:
- Local trial bundle:
kb/reports/cache/agentic-memory-refresh/2026-09-26-pilot/bundle. - Manifest SHA-256:
844f158f87906823765feb7371be95afc54929f3bdb6c28cb71aecb7c3d3bb82. - Matrix SHA-256:
d0901f96a8f43605f43e63a97a3f7a1eea21426c74244e96786ca3f17e2a7f95. - Population: Dynamic Cheatsheet, Mem0, Napkin; three code-grounded results, all with analysis cutoff 2026-09-26. Selection is explicit, not all-generated.
- Repository revision recorded by the bundle:
26fbe60caa5431b958a7f01569b404ffa1978bc9. This revision alone does not contain the newly created result bytes; the temporary bundle is the exact evidence for this workshop trial.
The complete three-result matrix builder, table renderer, and numerical analyzer all exited 0. The generated matrix equals the bundled matrix byte for byte. All 42 axis profiles were checked against their result frontmatter: values, assessment, evidence basis, and canonical-record references matched. The table records the same three run/source identities and all six input hashes.
The executable query ledger contains all fourteen dimension membership counts plus six focused queries, including exclusions and eligible run IDs. Re-executing its code produced identical output. The bounded synthesis trial used full bundled results and canonical records. Its author checked its own claims; no independent semantic review was commissioned. Current-source bundle verification passed immediately before the workshop trial was written.
Trial output directory: kb/reports/cache/agentic-memory-refresh/2026-09-26-pilot/.
It contains memory-systems.csv, memory-systems-table.md, statistics.txt,
query-output.jsonl, and the frozen bundle/. These ignored outputs support
the workshop trial; the historical and full-corpus public comparisons were not
replaced. The published per-system results are retained under their normal
skill-owned paths.
Acceptance decision and next batch
The end-to-end producer and consumer mechanics pass for these three cases. Statistical parity is conditional. Query Q3 (wired automatic writing and push) has zero eligible rows because every write-agency union has afforded basis; this is not evidence that no automatic writing exists. Mem0 and Dynamic Cheatsheet explicitly retain wired automatic routes alongside afforded manual input. Mem0's wired internal push is also hidden from the strong-value count when combined with afforded application pull. Lineage has the same aggregate coverage loss in all three rows.
The initial proposal (subsequently adopted as ADR 093) recorded the alternatives: keep complete-profile statistics with this limit, retain finer evidence, or define narrower comparison questions. Resolve or explicitly accept that limitation before bulk refresh. A changed contract requires new pilot runs and another acceptance check; do not patch these retained results.
After that decision, the next batch should test a document-grounded target, replacement of an existing generated review (Pond is already an observed unbounded-reader prerequisite), and enough additional code-grounded cases to reach the analyzer's five-pair redundancy threshold. Merely reaching five rows does not guarantee five eligible pairs. Broader source availability, every legacy boundary, and full-corpus completeness remain untested.
The three-system summary supports named mechanism contrasts. Representation counts require particular care: Mem0 includes model-derived dense encodings as parametric; Dynamic Cheatsheet classifies its imported numeric access vectors as symbolic with generation uninspected. These scoped labels must not be read as prevalence of vector use or model-weight training.
Producer observations
Napkin completed the independent memory pass, main-result integration, source verification, publication, and checked handoff. During integration, executable fallback paths contradicted an overstrong guarantee quoted from a source comment. The final result explicitly qualified the guarantee and retained the specialist disagreement in reconciliation. A matching quotation alone did not establish its attached claim.
Recoverable structural friction included adjacent canonical IDs separated by em dashes being interpreted as shorthand in Napkin, and ID-leading prose being interpreted as duplicate declarations in Mem0. Formatting was corrected while the runs remained running. These observations do not establish a need to change the schema or validator; no consumer or method code was changed for the pilot.