Evaluation

Type: types/tag-readme.md

This tag gathers how claims, changes, and agent outputs are judged and tested: oracles and LLM judges, what an evaluation warrants, the difference between a claim's warrant and its fit in a working theory, experiment, ablation, and benchmark design, and evidence records from Commonplace's own review and rewrite runs. No single defining note anchors it; warranted autonomy is bounded by oracle domain states the constraint most members work under: an evaluation licenses action only over the candidates its oracle can assess. Members span notes, evidence records, reference docs and proposals, and system analyses. Nearby but different: llm-reliability covers why LLM output deviates from intent and how to correct it; a note belongs here when its question is what a check, judge, or experiment can establish. This head is selective; use a scoped tag search for full membership.

Oracles and judges

Warrant and theory fit

Experiment and benchmark design

Review evidence

  • llm-reliability — why LLM output deviates and the correction machinery; oracle hardening sits on the boundary and many notes carry both tags
  • review-system — the shipped assay pipeline whose gates, verdicts, and runs several evidence records here measure
  • claims-and-grounding — whether a claim is supported by its sources; the grounding-verdict evidence records belong to both
  • learning-theory — the parent area for how systems learn; warrant and fit notes feed it
  • self-improving-systems — improvement loops need oracles and evidence of improvement; warranted autonomy and theory-fit notes are shared

Other tagged notes