Error messages that teach are a constraining technique

Type: kb/types/note.md · Tags: learning-theory, constraining

In agent systems, every error message the agent sees — linter output, test failures, hook warnings — is context that shapes its next action. Across production agent systems, error messages that teach the fix consistently produce more reliable behavior than bare failure signals. The error channel is an instruction channel.

This means the difference between FAIL and FAIL: description must be under 200 chars, yours is 247 — trim the last sentence is not cosmetic. Practitioners report the same pattern across harnesses: the richer message wins. The cost difference is negligible — same hook, better message. The reliability difference is large.

Lopopolo's report on OpenAI's Codex team puts it directly: "Linter error messages double as remediation instructions — every failure message teaches the agent the fix." And: "every mistake is a harness bug" — when an agent makes an error the system could have prevented through a better message, the system is at fault.

Orthogonal to enforcement strength

The constraining gradient moves from instructions through skills and hooks to scripts, trading flexibility for reliability. But there's a second axis: how much the enforcement artifact teaches when it fires. A blocking hook that says FAIL constrains maximally but informs minimally. A blocking hook that explains the fix constrains equally but informs maximally. Moving along this axis is cheap — it requires no change in trigger mechanism or enforcement strength, only better messages.

This is available at every layer. An instruction can say "check descriptions" or "descriptions must discriminate — if it paraphrases the title, rewrite it." A script can silently correct or log what it changed and why. The inform axis is orthogonal to the enforcement axis, and nearly free to improve.


Relevant Notes: