A consumption channel delivers force without the history that earned it

Type: kb/types/note.md · Tags: failure-modes, artifact-analysis, self-improving-systems

Behavioral authority attaches to a consumption path — a consumer, a channel, and a force — not to the bytes of an artifact. That precision has a converse consequence: whatever reaches the path receives its force unless the path checks whether the promotion is authorized. Here, earned means that an explicit authorization covers the content, its version, the role it will occupy, and the use being made of it. It does not mean that the content is guaranteed true.

This boundary makes a writable self-representation an attack surface rather than a maintenance chore. If retrieval is the wire a self-representation acts along, one failure occurs when search finds nothing and a represented constraint stays inert. Here, the wire works: it delivers content into a high-force position without a matching authorization transition.

Two ways in, one boundary

The adversarial case is familiar. Content reaches a channel that treats it as instruction: a poisoned rule file, a memory written from a compromised tool result, or text in retrieved material that reads as a directive. In an LLM context, instruction and data share one token medium; typed envelopes help only when the consumer preserves their distinction. This is the indirect-prompt-injection failure Greshake et al. identify in LLM-integrated applications, where an application “blurs the line between data and instructions.”

The innocent case has no attacker at all:

The runtime attack and the persistent false rule are different mechanisms with a shared boundary. The runtime case crosses from untrusted data into an instruction-bearing role during interpretation; the persistent case crosses through admission or lifecycle governance. In both cases, content occupies an operative position without an authorization that covers the transition. The cases therefore share a diagnostic question, not a complete control plan.

This distinction predicts incomplete defences. For persistent artifacts, controlling who may write without checking what was written admits authorized errors, while reviewing content without enforcing the reviewed entrance leaves a bypass. Ephemeral inputs additionally require consumption-side controls such as trust-tiered loading, instruction/data isolation, constrained capabilities, or a privilege quarantine.

Nor is it a wrong verdict

The boundary is also distinct from an evaluation gate that ran and erred. False-positive acceptance becomes operative describes an evaluator passing something it should have rejected. Its correction path is a stronger oracle — the mechanism that assesses the candidate — and warranted autonomy is bounded by what that oracle can assess. Those arguments assume the candidate went through the gate.

The failure here is that the channel does not necessarily check gate membership. Strengthening the oracle does nothing about a path that bypasses it. Conversely, a false positive may carry valid procedural authorization without carrying substantive truth. Authorization binding and oracle quality are therefore separate design problems, even though both can end in bad content acting.

Control must bind authorization to force

Controls can act at several distinct points along the boundary. They work together only when the system binds their state to the exact content consumed:

  • Runtime separation prevents untrusted content from silently becoming instruction or gaining privileged capabilities during a call.
  • Write authority restricts which paths may place persistent content in an operative position.
  • Review grants procedural authorization for a particular content version and scope; it does not establish infallibility.
  • Consumption enforcement and provenance make the channel verify that the authorization matches the version and use before conferring force. Provenance is evidence for that decision, not the decision itself.
  • Rollback withdraws future force after detection. It is recovery rather than prevention, and it remains a sovereignty capability because the owner must be able to restore or regenerate the artifact.

A channel closes the boundary only when every route to an operative position is subject to a mechanical rule that checks the relevant authorization state. A signed artifact, a trust-tiered loader, or a store whose sole write path binds approval to a content digest can supply that rule. Merely recording provenance or performing review somewhere upstream cannot.

This boundary belongs inside a reflective architecture: when operative artifacts are the self-representation, editing one modifies the system, so admission and consumption paths are part of the causal connection. Closing these paths is not free. Provenance can spend context, entrance controls spend throughput, and review spends evaluation capacity. Which controls are justified depends on the channel's force and deployment risk.

The cure states a claim of its own

Authorization metadata inherits the same problem. A verified flag or approved status is itself content in a trusted position, so a stale or unbacked mark can suppress the check that would have caught it. A derived copy of recomputable truth must therefore be checked or absent, while a non-recomputable attestation must be invalidated whenever its scope may no longer hold. This KB applies the latter rule to user-verified: a substantive edit strips the human attestation, and only another explicit act restores it.


Relevant Notes: