Evaluation
Type: types/tag-readme.md
This tag gathers how claims, changes, and agent outputs are judged and tested: oracles and LLM judges, what an evaluation warrants, the difference between a claim's warrant and its fit in a working theory, experiment, ablation, and benchmark design, and evidence records from Commonplace's own review and rewrite runs. No single defining note anchors it; warranted autonomy is bounded by oracle domain states the constraint most members work under: an evaluation licenses action only over the candidates its oracle can assess. Members span notes, evidence records, reference docs and proposals, and system analyses. Nearby but different: llm-reliability covers why LLM output deviates from intent and how to correct it; a note belongs here when its question is what a check, judge, or experiment can establish. This head is selective; use a scoped tag search for full membership.
Oracles and judges
- Warranted autonomy is bounded by oracle domain — autonomy is free, but warranted evaluation autonomy extends only as far as a confident oracle reaches
- Evaluation automation is phase-gated by comprehension — error analysis and demonstrated judge discrimination come before an automated loop can improve behavior rather than just score
- Weakly discriminated qualities tend to be underselected — conjecture: unequal oracle discrimination yields unequal enrichment under proposal selection
- Reasoning production is not reasoning evaluation — a critic can reconstruct the answer instead of checking the reasoning, so process validity needs its own check
- Review automation should target verifiable subroles before reviewer identity — decompose reviewer work into separately checkable parts before granting reviewer-level authority
- A search controller is tested by what it brings to stronger evaluation — provisional judgments route candidates to stronger checks; they are not acceptance claims
- Brainstorming: pairwise comparison and soft oracles — a staged plan for testing whether pairwise judges improve discrimination, stability, and calibration
Warrant and theory fit
- A claim's warrant does not determine its fit in a working theory — warrant and fit answer different questions; either can hold without the other
- System use provides evidence of theory fit, not independent warrant — live use tests integration and usefulness; factual, source, or scope evidence is still needed for warrant
- System use is an initial selection environment when theory fit lacks a fixed oracle — distributed consequences of use can select claims where no fixed oracle exists
- Disconnected witnesses do not establish a causal path through theory — evidence of learning through theory must join one causal path
- Mixed epistemic status must be preserved below the document level — observations, deductions, and plausible explanations keep their separate warrant inside one document
Experiment and benchmark design
- An experiment identifies only the contrast it actually runs — missing comparisons and bundle-to-component attribution overstate causal conclusions
- A retained-theory intervention isolates one surface — intervening on retained theory estimates that surface's contribution, not whole-program theory possession
- A benchmark that holds the client fixed exports the least-warrantable decisions — fixed-client benchmarks measure worker capability and leave closure untested
- Prompt ablation converts human insight into deployable framing — controlled variation against a human-verified finding picks the framing agents can execute
- Systematic prompt variation serves verification and diagnosis — varying prompts decorrelates checks or measures brittleness; it does not test explanatory-reach
- Ablation baselines for the declared objective — proposal: matched repository tasks with curated theory, review, or episodes removed
- Trajectory-aware evaluation of transforming agent workflows — proposal: compare blinded output-only and trajectory-aware judges before adding trajectory checks
Review evidence
- Full improvement pass closure — how the shipped workflow reassays final note bytes and stops without claiming convergence
- A five-link cap missed four grounding findings in twelve reviews — paired assay: fuller reading of linked artifacts surfaced findings the cap hid
- An independent pass tightened three of four Pirolli grounding verdicts — separating source reconstruction from claim judgment changed verdicts; a candidate control, not a proven cause
- Single-artifact review bundles still cut Claude costs substantially — cache-weighted telemetry for the single-artifact bundle refactor
- Three simplification passes exposed different clarity–precision tradeoffs — broad style guidance, a compact cue, and exhaustive local review compared on one article
Related Tags
- llm-reliability — why LLM output deviates and the correction machinery; oracle hardening sits on the boundary and many notes carry both tags
- review-system — the shipped assay pipeline whose gates, verdicts, and runs several evidence records here measure
- claims-and-grounding — whether a claim is supported by its sources; the grounding-verdict evidence records belong to both
- learning-theory — the parent area for how systems learn; warrant and fit notes feed it
- self-improving-systems — improvement loops need oracles and evidence of improvement; warranted autonomy and theory-fit notes are shared
Other tagged notes
- A better-factory claim compares operative states under an antecedent assessment relation - The improvement claim's relata are predecessor and operative-successor states and its relation is declared before the development it judges; evaluator location is a separate declaration from the learner boundary
- A checked outcome licenses retaining an episode, not abstracting its explanation - One result-only check can warrant retaining an episode as evidence, but abstracting its explanation also needs evidence about a faithful producing process and an explicit scope boundary
- A claim without external assessment carries three obligations - Without external assessment a claim needs its own contradiction-and-support rule, a comparison level for objective change, and a performance measure it does not grade itself, plus attribution when it asserts a cause
- A context-operation interface bounds the projections its policy can realize - Explains why improving context selection within a fixed operation interface cannot establish that the interface admits every useful active-context projection.
- A linked note discharges its own grounding, so a citing note owes representation, not re-grounding - A cited source imposes a grounding obligation; a claim-titled note that already passed its own grounding review imposes only a representation obligation — with the preconditions that keep the distinction and why it is not a paraphrase ledger
- A note is an atomic step relative to the check that reads it - Two independent bounds on a note: one claim sized to the reader's bounded context, and one checkable inference sized to the checker's single pass — for the grounding check the unit is the unquoted source
- A proximate target is checked for achievement, not for warrant - Between an improvement objective and its oracles sits a target level — a property pursued because it is held to serve the objective — whose linking claim no check in the loop evaluates
- A quotes-route rollout grounded more claim uses without earning claim identifiers - A Commonplace grounding rollout recorded 30% grounded claim uses under a paraphrased claims ledger and 75% under verbatim quotes or pinned snapshots, with no case needing claim identifiers; the non-random cohorts make the gap descriptive
- A vibe-noting trace shows persistence enables revision, not certification - Evidence from one Commonplace note history: persistence enabled later semantic development while review exposed omitted risks, attribution drift, a link error, and an unresolved authority boundary
- Academic Research Skills - Academic Research Skills as a prompt-defined Claude Code research pipeline with narrow executable checks, host-dependent orchestration, protocol-only resume, and conflicting terminal gate rules
- Activate Behavior-Changing Memory Before The Mistake - Behavior-changing memory must activate before relevant actions rather than waiting for explicit retrospective search
- AI Agents in Depth - Whole-book comparison of AI Agents in Depth with Commonplace, separating broad architectural convergence from differences in memory admission, epistemic warrant, governance, and orchestration
- An accepted edit verifies the change, not the rule - Human acceptance of an edit is a strong oracle for 'this change was wanted here' but a weak oracle for 'this generalizes' — mining rules from accepted edits inherits instance-level verification while the generalization step stays oracle-poor
- An adversarial human-agent loop can reconstruct the writing-is-thinking filter - The writing-is-thinking filter is the loop's, not the pen's — an adversarial human-agent loop can reconstruct what naive delegation loses, but only while the human stays the judge
- Automated synthesis is missing good oracles - Generating synthesis candidates (cross-note connections, novel combinations) is easy — LLMs do it readily. The hard part is evaluating whether a candidate is genuine insight or noise.
- Automated tests for text - Text artifacts can be tested with the same pyramid as software — deterministic checks, LLM rubrics, corpus compatibility — built from real failures not taxonomy
- Brainstorming: maintainability oracles for agentic development - Explores candidate signals, calibration experiments, authority levels, and workflow placements for evaluating maintainability in agent-generated code
- Candidacy evidence licenses escalation to assessment, not acceptance - Separates candidacy evidence, which routes a hypothesis to costly assessment, from verdict evidence, which decides it; pricing and source-grounding cases provide two worked witnesses
- Causal and proof obligations are two formal routes to assessing explanatory-reach - Causal and proof obligations demonstrate two ways formal symbolic systems can assess explanatory-reach inside a warranted model
- Cheap generation breaks text volume as an effort signal - When text is cheap to expand but costly to verify, length stops evidencing author effort and can instead warn that the reviewer inherits unperformed checking
- Competing causal theories can guide distinguishing experiments - Why observationally equivalent mechanisms can tell a theory builder what to test next: a noisy binary example separates evidence acquisition from choosing or verifying an explanation.
- Compounding is tested in later improvement, not by the accepting metric - Compounding evidence must come from later improvement episodes through displaced productivity measures and causal traces, not from the metric that accepted the earlier change
- Elicitation requires maintained question-generation systems - Four elicitation strategies ordered by user expertise required, composable into review architectures with maintenance loops that prevent ossification
- Evaluate Memory By Effects, Not By Existence - Memory should be evaluated by downstream effects on tasks, artifacts, answers, behavior, context efficiency, and lineage alignment
- Exact implementation does not validate a requirement against its objective - An artifact can exactly implement a requirement while the requirement remains a conjectured proxy for a declared objective; assess each named path separately, and attribute failure to the link without erasing local correctness
- False-positive generation is filtered; false-positive acceptance becomes operative - False-positive generation faces evaluation before retention, while false-positive acceptance becomes operative and can compound
- Feedback-trained memory management is oracle-dependent even when its operations are hand-designed - Fixed and merely runtime-responsive memory rules need no training oracle; outcome-driven updates do, while noisy rankings weaken learning and misaligned ones teach the wrong ordering
- Full write briefs cut edit drift; one-line briefs did not - Pre-registered pilot, 95 edit runs on 10 KB documents: a full retained write brief cut dropped commission items by about two thirds, a brief rebuilt from the commissioned document did nearly as well, and a one-line brief matched no brief
- Generation confidence does not by itself certify soundness - Distinguishes next-token probability from factual truth and inferential validity: confidence can support correctness decisions only after task-specific validation, and high-assurance acceptance still needs a separate check
- In one episode, recognition appeared only in the corpus-loaded run - One 2026-09-01 episode: a repository-free synthesis re-derived retained notes and reproposed rejected framings while the corpus-loaded session recognized them; an uncontrolled bundle, recorded as a starting point for better contrasts
- Knowledge storage does not imply contextual activation - Separates knowledge that exists, knowledge loaded into context (read-back), and knowledge that actually changes behavior (activation); explains why retrieval and long context do not guarantee activation
- Knowledge-access architecture must be evaluated end to end, not by retrieval alone - Explains why retrieval measures and storage-substrate labels cannot proxy for task-relative quality across discovery, loading, transformation, activation, and upkeep
- Known-target discovery benchmarks show reachability, not discovery closure - Distinguishes backcast and reinvention benchmarks from autonomous discovery: they show that target insights are reachable from supplied ingredients, not that a system can select and verify new discoveries prospectively.
- Learning inside a fixed decomposition inherits its mistakes - Why optimization cannot repair consequential distinctions, responses, or mappings outside the effective update space of a fixed task decomposition
- LLM output deviation requires three-way diagnosis because remedies target different relations - For a fixed assembled input, whether V exceeds I, whether D escapes V, and how D's spread affects realization are three diagnostic questions with different primary repair surfaces
- Memory-backed personalization can look like model improvement - Distinguishes user-specific gains supplied by retained intent from gains in the model that interprets the assembled context.
- Open-ended improvement must allocate search before decisive evaluation is available - Open-ended improvement must choose which questions, candidates, experiments, or proof paths to develop before decisive evidence about them is available; even a Gödel machine's proof gate retains this prior search problem
- Oracle accumulation improves selection for later candidates in its maintained domain - A failure retained as a lesson helps tasks that retrieve it; retained as a maintained check it improves selection for later candidates in its domain and amortizes validation
- Oracle strength spectrum - Exploratory framework — oracle strength, how cheaply correctness can be verified, as the gradient underlying the exact-spec/proxy-theory distinction, with an oracle-hardening pipeline
- Process structure and output structure are independent levers - Distinguishes constraints on reasoning steps from constraints on result shape and identifies the evidence needed to separate their effects
- Psychology-to-agent transfer needs per-principle failure-mode testing - Brainstorming a methodology for evaluating cognitive-science-to-agent transfer — assembled from three existing KB notes and tested against Youssef's five psychology principles as worked examples
- Quality signals for KB evaluation - Catalogues graph-topology, content-proxy, and LLM-hybrid signals that could be combined into a weak composite oracle to drive a mutation-based KB learning loop without requiring usage data.
- Reach-assessment - Definition — judging whether a commitment's claimed explanatory-reach is genuine across natural-language, symbolic, and distributed-parametric forms
- Reliability dimensions map to oracle-hardening stages - The four reliability dimensions from Rabanser et al. (consistency, robustness, predictability, safety) each harden a different oracle question — mapping empirical agent evaluation onto the oracle-strength spectrum
- Revision guided by rationale needs faithfulness, not just legibility - When revision of an addressable theory relies on rationale to locate a failed premise, misleading rationale can direct repair to the wrong part; rationale is one optional repair aid
- Selecting an LLM output fixes a result, not its interpretation - Selecting one LLM output for operative reuse creates a stable artifact-testing target without resolving ambiguity inside the text, so generator and artifact tests answer different questions
- Structured-prompt gains do not establish training-distribution selection - Formatting compliance, extra computation, and task decomposition can mimic distribution-selection gains, so prompt performance alone cannot identify the mechanism
- Task families and product families classify different things - Task families group obligations or evaluations; software product families group products through declared commonality, variability, and reusable production scope
- Technical constraints turn KB objective-function choice from philosophy into engineering - Three technical constraints and the codification lever make KB objective-function choice testable engineering, not philosophy; goals set the loss, local contracts specialize it, and oracle strength differs by objective
- The augmentation-automation boundary is discrimination not accuracy - Crossing from augmentation to automation requires per-instance discrimination, not aggregate accuracy — discrimination is empirically stagnant, so scaling capability alone cannot cross the boundary
- The boundary of automation is the boundary of verification - Synthesis — oracle theory, labor economics, frontier-lab capability predictions, and supply-chain integrity evidence converge on verification cost as the primary structural determinant of automation
- The verifiability gradient - Symbolic artifacts sit on a gradient from loose natural-language to deterministic code; higher-verifiability artifacts support tighter iteration loops, and learning moves artifacts along it in both directions
- Theory warrant should be tracked at the finest granularity evidence licenses - Treat support for a theory as warrant for only the most specific claim, conjunction, model, and scope the evidence identifies; do not distribute joint warrant beyond what it entails without additional attribution
- Two rewrites exposed a syntax-or-repetition tradeoff - Evidence from two ASD-STE100-inspired passes over one note: unguarded sentence splitting lost semantic relations, while guarded splitting preserved them by adding 4.9% more words
- Universal software factory needs a declared universality axis - Universal software factory is ambiguous unless the universality axis, covered class, supplied inputs, adequacy relation, and resource bounds are declared
- Use tests a decomposition locally; retained rationale is what makes transfer testable - Running a decomposition confirms only that it sufficed here; because many force-sets fit the same split, rationale retained at design time is what gives a transfer claim an antecedent to test
- Warranted transfer out of the human cut leaves people the hardest-to-warrant decisions - When a system preferentially transfers decisions whose premises, criteria, and checks are available, the remaining human decisions become harder to warrant per decision; this predicts a residue composition, not structural computational openness