Evaluation
Type: types/tag-readme.md
This tag gathers how claims, changes, and agent outputs are judged and tested: oracles and LLM judges, what an evaluation warrants, the difference between a claim's warrant and its fit in a working theory, experiment, ablation, and benchmark design, and evidence records from Commonplace's own review and rewrite runs. No single defining note anchors it; warranted autonomy is bounded by oracle domain states the constraint most members work under: an evaluation licenses action only over the candidates its oracle can assess. Members span notes, evidence records, reference docs and proposals, and system analyses. Nearby but different: llm-reliability covers why LLM output deviates from intent and how to correct it; a note belongs here when its question is what a check, judge, or experiment can establish. This head is selective; use a scoped tag search for full membership.
Oracles and judges
- Warranted autonomy is bounded by oracle domain — autonomy is free, but warranted evaluation autonomy extends only as far as a confident oracle reaches
- Evaluation automation is phase-gated by comprehension — error analysis and demonstrated judge discrimination come before an automated loop can improve behavior rather than just score
- Weakly discriminated qualities tend to be underselected — conjecture: unequal oracle discrimination yields unequal enrichment under proposal selection
- Reasoning production is not reasoning evaluation — a critic can reconstruct the answer instead of checking the reasoning, so process validity needs its own check
- Review automation should target verifiable subroles before reviewer identity — decompose reviewer work into separately checkable parts before granting reviewer-level authority
- A search controller is tested by what it brings to stronger evaluation — provisional judgments route candidates to stronger checks; they are not acceptance claims
- Brainstorming: pairwise comparison and soft oracles — a staged plan for testing whether pairwise judges improve discrimination, stability, and calibration
Warrant and theory fit
- A claim's warrant does not determine its fit in a working theory — warrant and fit answer different questions; either can hold without the other
- System use provides evidence of theory fit, not independent warrant — live use tests integration and usefulness; factual, source, or scope evidence is still needed for warrant
- System use is an initial selection environment when theory fit lacks a fixed oracle — distributed consequences of use can select claims where no fixed oracle exists
- Disconnected witnesses do not establish a causal path through theory — evidence of learning through theory must join one causal path
- Mixed epistemic status must be preserved below the document level — observations, deductions, and plausible explanations keep their separate warrant inside one document
Experiment and benchmark design
- An experiment identifies only the contrast it actually runs — missing comparisons and bundle-to-component attribution overstate causal conclusions
- A retained-theory intervention isolates one surface — intervening on retained theory estimates that surface's contribution, not whole-program theory possession
- A benchmark that holds the client fixed exports the least-warrantable decisions — fixed-client benchmarks measure worker capability and leave closure untested
- Prompt ablation converts human insight into deployable framing — controlled variation against a human-verified finding picks the framing agents can execute
- Systematic prompt variation serves verification and diagnosis — varying prompts decorrelates checks or measures brittleness; it does not test explanatory-reach
- Ablation baselines for the declared objective — proposal: matched repository tasks with curated theory, review, or episodes removed
- Trajectory-aware evaluation of transforming agent workflows — proposal: compare blinded output-only and trajectory-aware judges before adding trajectory checks
Review evidence
- Full improvement pass closure — how the shipped workflow reassays final note bytes and stops without claiming convergence
- A five-link cap missed four grounding findings in twelve reviews — paired assay: fuller reading of linked artifacts surfaced findings the cap hid
- An independent pass tightened three of four Pirolli grounding verdicts — separating source reconstruction from claim judgment changed verdicts; a candidate control, not a proven cause
- Single-artifact review bundles still cut Claude costs substantially — cache-weighted telemetry for the single-artifact bundle refactor
- Three simplification passes exposed different clarity–precision tradeoffs — broad style guidance, a compact cue, and exhaustive local review compared on one article
Related Tags
- llm-reliability — why LLM output deviates and the correction machinery; oracle hardening sits on the boundary and many notes carry both tags
- review-system — the shipped assay pipeline whose gates, verdicts, and runs several evidence records here measure
- claims-and-grounding — whether a claim is supported by its sources; the grounding-verdict evidence records belong to both
- learning-theory — the parent area for how systems learn; warrant and fit notes feed it
- self-improving-systems — improvement loops need oracles and evidence of improvement; warranted autonomy and theory-fit notes are shared
Other tagged notes
- A linked note discharges its own grounding, so a citing note owes representation, not re-grounding - A cited source imposes a grounding obligation; a claim-titled note that already passed its own grounding review imposes only a representation obligation — with the preconditions that keep the distinction and why it is not a paraphrase ledger
- A note is an atomic step relative to the check that reads it - Two independent bounds on a note: one claim sized to the reader's bounded context, and one checkable inference sized to the checker's single pass — for the grounding check the unit is the unquoted source
- A vibe-noting trace shows persistence enables revision, not certification - Evidence from one Commonplace note history: persistence enabled later semantic development while review exposed omitted risks, attribution drift, a link error, and an unresolved authority boundary
- Academic Research Skills - Academic Research Skills as a prompt-defined Claude Code research pipeline with narrow executable checks, host-dependent orchestration, protocol-only resume, and conflicting terminal gate rules
- AI Agents in Depth - Whole-book comparison of AI Agents in Depth with Commonplace, separating broad architectural convergence from differences in memory admission, epistemic warrant, governance, and orchestration
- Brainstorming: maintainability oracles for agentic development - Explores candidate signals, calibration experiments, authority levels, and workflow placements for evaluating maintainability in agent-generated code
- Elicitation requires maintained question-generation systems - Four elicitation strategies ordered by user expertise required, composable into review architectures with maintenance loops that prevent ossification
- Full write briefs cut edit drift; one-line briefs did not - Pre-registered pilot, 95 edit runs on 10 KB documents: a full retained write brief cut dropped commission items by about two thirds, a brief rebuilt from the commissioned document did nearly as well, and a one-line brief matched no brief
- In one episode, recognition appeared only in the corpus-loaded run - One 2026-09-01 episode: a repository-free synthesis re-derived retained notes and reproposed rejected framings while the corpus-loaded session recognized them; an uncontrolled bundle, recorded as a starting point for better contrasts
- Knowledge storage does not imply contextual activation - Separates knowledge that exists, knowledge loaded into context (read-back), and knowledge that actually changes behavior (activation); explains why retrieval and long context do not guarantee activation
- Psychology-to-agent transfer needs per-principle failure-mode testing - Brainstorming a methodology for evaluating cognitive-science-to-agent transfer — assembled from three existing KB notes and tested against Youssef's five psychology principles as worked examples
- Revision guided by rationale needs faithfulness, not just legibility - When revision of an addressable theory relies on rationale to locate a failed premise, misleading rationale can direct repair to the wrong part; rationale is one optional repair aid
- Two rewrites exposed a syntax-or-repetition tradeoff - Evidence from two ASD-STE100-inspired passes over one note: unguarded sentence splitting lost semantic relations, while guarded splitting preserved them by adding 4.9% more words