Evaluation
Type: kb/types/tag-readme.md
What works, what doesn't, what needs testing. Empirical observations about KB operations, prompt design, and techniques from other systems.
Notes
- cludebot — techniques from cludebot worth borrowing; richest trajectory-to-lesson loop reviewed
- prompt-ablation-converts-human-insight-to-deployable-framing — methodology for testing prompt framings
- single-artifact review bundles still cut Claude costs substantially after cache-aware weighting — cache-weighted April 2-4, 2026 evidence that the single-artifact refactor remained a significant cost win under Anthropic's prompt-caching prices
- systematic-prompt-variation-serves-verification-and-diagnosis-not-explanatory-reach-testing — controlled variation as a family of methods: decorrelating checks, measuring brittleness, and distinguishing both from Deutsch-style explanatory-reach review
- brainstorming-how-to-test-whether-pairwise-comparison-can-harden-soft-oracles — experimental ladder for comparing scalar and pairwise judges before treating pairwise ranking as a stronger soft oracle
Other tagged notes
- A benchmark that holds the client fixed exports the least-warrantable decisions by design - A fixed-client benchmark measures worker capability; it leaves broader closure untested when the client supplies internal production decisions, while ordinary user requirements and acceptance may remain external
- A claim's warrant does not determine its fit in a working theory - Independent warrant and fit in a working theory answer different questions: a warranted claim may fit poorly, while apparent fit may be produced by an unwarranted or already-assumed claim
- A five-link cap missed four grounding findings in twelve reviews - A paired Commonplace assay found five capped-versus-uncapped grounding outcome divergences; one reproduced as reviewer noise, while four appeared only after fuller reading reached 6–16 linked artifacts
- A linked note discharges its own grounding, so a citing note owes representation, not re-grounding - A cited source imposes a grounding obligation; a claim-titled note that already passed its own grounding review imposes only a representation obligation — with the preconditions that keep the distinction and why it is not a paraphrase ledger
- A note is an atomic step relative to the check that reads it - Two independent bounds on a note: one claim sized to the reader's bounded context, and one checkable inference sized to the checker's single pass — for the grounding check the unit is the unquoted source
- A retained-theory intervention isolates one surface, not the whole program theory - An intervention on retained theory estimates that surface's causal contribution under matched conditions; influence, explanatory guidance, acquisition, and whole-system theory possession remain different claims
- A search controller is tested by what it brings to stronger evaluation - A search controller should be evaluated by the branches and probes it routes into stronger evaluation, not by treating every provisional judgment as an acceptance claim
- A vibe-noting trace shows persistence enables revision, not certification - Evidence from one Commonplace note history: persistence enabled later semantic development while review exposed omitted risks, attribution drift, a link error, and an unresolved authority boundary
- An independent pass tightened three of four Pirolli grounding verdicts - A Commonplace grounding case changed three of four support verdicts after separating source reconstruction from target-claim judgment, making bilateral isolation a candidate control rather than a proven cause.
- Brainstorming: maintainability oracles for agentic development - Explores candidate signals, calibration experiments, authority levels, and workflow placements for evaluating maintainability in agent-generated code
- Disconnected witnesses do not establish a full causal path through theory - Theory use, outcome, theory revision, and later use establish theory-mediated learning only when their witnesses identify the joins of the same full causal path
- Elicitation requires maintained question-generation systems - Four elicitation strategies ordered by user expertise required, composable into review architectures with maintenance loops that prevent ossification
- Evaluation automation is phase-gated by comprehension - Optimization loops need diagnostic error analysis and demonstrated judge discrimination before automation can improve behavior rather than just score
- In one episode, recognition appeared only in the corpus-loaded run - One 2026-09-01 episode: a repository-free synthesis re-derived retained notes and reproposed rejected framings while the corpus-loaded session recognized them; an uncontrolled bundle, recorded as a starting point for better contrasts
- Knowledge storage does not imply contextual activation - Separates knowledge that exists, knowledge loaded into context (read-back), and knowledge that actually changes behavior (activation); explains why retrieval and long context do not guarantee activation
- Mixed epistemic status must be preserved below the document level - A document can combine observations, deductions, and plausible explanations; KB writing and review must retain which claims and transitions have which warrant.
- Reasoning production is not reasoning evaluation - Review and critique systems need independent process-validity checks because a model can substitute answer reconstruction for reasoning evaluation
- Review automation should target verifiable subroles before reviewer identity - Scholarly-review automation should decompose reviewer work into separately verifiable subroles before giving an AI system reviewer-level authority
- Revision guided by rationale needs faithfulness, not just legibility - When revision relies on a rationale to locate a failed premise, a misleading rationale can direct repair to the wrong part; retained rationale is optional for theory refinement
- System use is an initial selection environment when theory fit lacks a fixed oracle - When no complete fixed oracle decides whether a claim belongs in a working theory, distributed consequences of live system use can provide an initial selection environment
- System use provides evidence of theory fit and causal usefulness, not independent warrant - Consequences of using a claim in a live system can test its integration and causal usefulness, but independent factual, formal, source, or scope evidence is still needed for its warrant
- Three simplification passes exposed different clarity–precision tradeoffs - Evidence from three independent rewrites of one mature article: broad style guidance ranked best overall, a compact style cue improved rhythm but drifted, and exhaustive local review barely changed the text
- Two rewrites exposed a syntax-or-repetition tradeoff - Evidence from two ASD-STE100-inspired passes over one note: unguarded sentence splitting lost semantic relations, while guarded splitting preserved them by adding 4.9% more words
- Warranted autonomy is bounded by oracle domain - Bare autonomy is free, but warranted evaluation autonomy extends only to the candidates an oracle can assess with the required confidence
- Weakly discriminated qualities tend to be underselected - Statistical conjecture: under named proposal-selection conditions, unequal oracle discrimination yields unequal enrichment; absolute degradation needs an additional directional mechanism