Error messages that teach are a constraining technique

Type: kb/types/note.md · Tags: learning-theory, constraining

In agent systems, every error message the agent sees — linter output, test failures, hook warnings — is context that shapes its next action. The error channel is an instruction channel.

This means the difference between FAIL and FAIL: description must be under 200 chars, yours is 247 — trim the last sentence is not cosmetic. The cost difference is negligible — same hook, better message. The reliability difference is large.

Lopopolo's report on OpenAI's Codex team puts it directly: "Linter error messages double as remediation instructions — every failure message teaches the agent the fix." And: "every mistake is a harness bug" — when an agent makes an error the system could have prevented through a better message, the system is at fault.

Orthogonal to enforcement strength

The constraining gradient moves from instructions through skills and hooks to scripts, trading flexibility for reliability. But there's a second axis: how much the enforcement artifact teaches when it fires. A blocking hook that says FAIL constrains maximally but informs minimally. A blocking hook that explains the fix constrains equally but informs maximally. Moving along this axis is cheap — it requires no change in trigger mechanism or enforcement strength, only better messages.

This is available at every layer. An instruction can say "check descriptions" or "descriptions must discriminate — if it paraphrases the title, rewrite it." A script can silently correct or log what it changed and why. The inform axis is orthogonal to the enforcement axis, and nearly free to improve.


Relevant Notes: