Failure modes

Type: types/tag-readme.md

This tag gathers recurring failure modes of an agent-operated KB: ways knowledge and claims fail to do their job even though they exist on disk or have passed review. The anchor is knowledge storage does not imply contextual activation, the distinction between knowledge existing, being loaded, and affecting behavior. The members cover four kinds of failure: activation failures (stored knowledge not discovered, loaded, or acted on), authority failures (content gaining binding force it was never granted), claim-repair escapes (a claim survives review by becoming vaguer, analytic, or immune to refutation), and lessons generalized past their evidence. Nearby but different: llm-reliability covers how a model deviates from instructions and evidence and how those deviations are corrected; a failure belongs here when the fault lies in how the KB stores, delivers, or states knowledge.

Activation and delivery

Claim-repair escapes

Over-generalization and false assurance

  • LLM reliability — adjacent area: the deviation taxonomy and correction machinery for failures in how models interpret instructions and evidence
  • Evaluation — methods for detecting whether failures are real and whether interventions improve behavior