Context engineering
Type: types/tag-readme.md
This tag gathers work on getting the right knowledge into a bounded LLM context at the right time: routing and retrieval, loading and prompt assembly, scoping, context budgets and degradation, scheduling work across context windows, and deciding what to store so it can be loaded later. The defining note is context engineering, which names routing, loading, scoping, and maintenance as the operational core. Members span kb/notes/, reference proposals for write-context assembly, and system analyses in kb/agentic-systems/. Nearby but different: agent-memory covers what an agent retains across sessions and under what authority; a note belongs here when its question is how knowledge reaches a bounded call, whether or not it was retained as memory. This head is selective; use a scoped tag search for full membership.
Core Claims
- Context engineering - the definition: the discipline of designing systems around bounded-context calls, with routing, loading, scoping, and maintenance as its core
- Context efficiency is the central design concern in agent systems - why context is the scarce resource, and why per-window degradation binds before token cost
- Design for the first-time human, except on access cost - explains why human-facing materializations and agent-facing query paths can share a source of truth while following different access modes
- semantic sub-goals that exceed one context window become scheduling problems - explains when context limits force orchestration instead of a single larger prompt
- A context-operation interface bounds the projections its policy can realize - separates structural projection reach from policy success within one interface
- Addressability grain, not compression ratio, sets a matched selective-read floor - isolates the retrieval-floor condition for a known question that maps to one discriminating unit on each path, while leaving fan-out, sufficiency, reliability, and net value separate
- Opposed recompute factors do not decide documentation segmentation - why crossed savings and recurrence rankings do not order cache value, and why cache value alone does not decide whether a second content layer pays
- A derived copy of recomputable truth must be checked or absent - names when a recomputable value is safe to inline for context economy: only when a validator can re-derive and check it, otherwise it must stay a live read
Methodology Activation
- Capable agents need methodology selection - relevant but incompatible approaches force choosing a governing methodology, not just supplying knowledge
- Weight-resident methodologies compress behavior in context - a compact cue can activate a methodology already in the weights, at the cost of exact specification
Write-Context Assembly
- Deterministic write-context assembly - proposal: code assembles a target's fixed authoring context with closed input roles
- ADR 092: Write briefs are optional sidecars named by a validated pointer - decision: keep an artifact-specific commission beside its document and deliver it to writers
Related Tags
- Agent memory - the retention side: what agents keep across sessions and how it is activated; most memory-requirements notes carry both tags
- Computational model - the bounded-call substrate context engineering operates on, including framework-owned loops and when scheduling must become explicit
- Deploy-time learning - the post-release change that context machinery helps absorb without retraining
Other tagged notes
- A compact, refreshable whole-picture narrative can replace infeasible fragment reconciliation - Holistic rewrite shifts reconciliation from each consumer to the author, but only when the whole-picture narrative can fit within effective context and be refreshed before the narrative goes stale
- A knowledge base should support fluid resolution-switching - Defines resolution-switching as movement among KB views with different scope and detail, then inventories the mechanisms and limits of that qualitative criterion
- A linked note's durable payload is what its consumption path cannot reliably supply - Retain the recognition anchor and rationale the intended consumption path cannot reliably supply — an enforced path can carry the anchor itself; reconstructable framework recap factors into the linked artifact, tested by downstream effects
- A retrieval miss is a local reflective-path failure - A missed relevant artifact leaves its represented aspect inert for the affected task and discovery route, while other loading paths and reflective aspects can remain causally connected
- A specific intent may out-yield local rationales, but contingent facts stay separate - Conjectures that an unrecoverable governing intent yields more local rationale per token than rationale snippets, while contingent design facts need their own record
- Academic Research Skills - Academic Research Skills as a prompt-defined Claude Code research pipeline with narrow executable checks, host-dependent orchestration, protocol-only resume, and conflicting terminal gate rules
- Access burden and transformation burden are distinct query dimensions - Separates the system-relative cost of finding required inputs from producing an answer, so query systems can diagnose which work remains as retrieval and reasoning interact
- Activate Behavior-Changing Memory Before The Mistake - Behavior-changing memory must activate before relevant actions rather than waiting for explicit retrospective search
- Active work state is not retrospective memory or chat history - Active work state needs current pointers, evidence gates, and closure; treating it as retrospective memory or chat history preserves the wrong state
- Adaptation signals choose pressure; artifact analysis chooses the retained surface - Maps agentic-adaptation signals onto artifact-analysis axes so KB learning records which retained surface changes, what authority it gains, and how to review it
- Agent memory is a crosscutting concern, not a separable niche - Memory decomposes into storage (solved), retrieval/activation (context engineering), and learning (learning theory) — treating it as a standalone category hides that the hard problems are at the intersections
- Agent memory needs discoverable, loadable, composable, trusted knowledge under bounded context - Distinguishes four use-time requirements for remembered knowledge—discoverability, loadability, composability, and calibrated trust—from system-level activation.
- Agent Memory Requirements - Navigation hub for concrete agent-memory requirements extracted from the memory-system design synthesis
- Agent-runtime analysis should separate scheduling, context assembly, and external state - For runtimes composed of bounded model calls, separating control progression, per-call context, and external state or action services localizes failures even when one implementation owns all three
- Agents navigate by deciding what to read next - Models agent navigation as repeated follow/skip judgment under bounded context: cue diagnosticity must repay its own context cost, so longer pointer context is not automatically better
- AGENTS.md should be organized as a control plane - Theory for deciding what belongs in AGENTS.md using loading frequency and failure cost, with layers, exclusion rules, and migration paths
- AI Agents in Depth - Whole-book comparison of AI Agents in Depth with Commonplace, separating broad architectural convergence from differences in memory admission, epistemic warrant, governance, and orchestration
- Always-loaded context mechanisms in agent harnesses - Survey of always-loaded context mechanisms across agent harnesses — system prompt files, capability descriptions, memory, and configuration injection — cataloguing what each carries, how write policies differ, and where the gaps are
- An insufficient summary precedes the source rather than replacing it - When a summary cannot license a reliability-compliant stop, the authoritative fallback remains in the path; only fallback work the summary removes can offset its own cost
- Artifact function as a routing field - Proposal: whether an artifact_function declaration should expose a document's intended whole-artifact job for writing and review routing without asserting atomicity
- beads_rust - beads_rust as a local active-work and coordination substrate: transactional CLI claims and workflow gates, explicit external execution and Git boundaries, and weaker parity across MCP, inherited-context, and shipped instruction paths
- Borrowing can operate through retained artifacts or weight activation - Established external methodologies can become operative either by being explicitly retained in the system or by activating a model's pretrained representation; the two routes trade context economy against inspectability and revisability
- Checked inline blocks for shared instruction text - Proposal: reuse natural-language authoring mechanics in specialized writer prompts through literal inlining backed by deterministic source-to-copy checks
- Context contamination operates below an agent's compliance reasoning - A controlled test found fine-grained stance drift despite explicit detection and refusal; exclusion guarantees non-exposure, while instruction-level mitigation remains an empirical question
- Conversation vs prompt refinement in agent-to-agent coordination - Conversation preserves the execution trace; prompt refinement compresses it into a clean handoff. The right choice depends on architecture and how much intermediate work should survive
- Create Memory Directly - Direct memory creation preserves live understanding by writing useful artifacts before later trace extraction loses structure
- Decomposition heuristics for bounded-context scheduling - Working heuristics for symbolic scheduling over bounded LLM calls — separate selection from joint reasoning, choose representations not just subsets, save reusable intermediates in scheduler state
- Design rationale must preserve decision premises its interpreter cannot regenerate - Retention test for source-checkout design rationale: keep current decision premises not faithfully recoverable from implementation, git, and general knowledge; treat recoverable, role-free explanation as a cache
- Designing a Memory System for LLM-Based Agents - Derives agent-memory design pressures and links to a requirements inventory for agents designing or evaluating memory systems
- Elicitation requires maintained question-generation systems - Four elicitation strategies ordered by user expertise required, composable into review architectures with maintenance loops that prevent ossification
- Evaluate Memory By Effects, Not By Existence - Memory should be evaluated by downstream effects on tasks, artifacts, answers, behavior, context efficiency, and lineage alignment
- Flat memory predicts specific cross-contamination failures that are empirically testable - Flat memory predicts three cross-contamination failures — search pollution, identity scatter, insight trapping — testable via an observation protocol against real agent systems
- Frontloading spares execution context - Pre-computing known instruction inputs and inserting their results spares execution-context budget inside a later LLM call
- History has one chance to become checkable - An artifact's production history is convertible to later-checkable form only at production time, via records/attestation or re-derivability; after that a bounded reviewer sees only carried state
- Import External Knowledge Into Internal Form - Agent memory systems need import paths when authoritative project knowledge already exists outside the memory substrate
- In one episode, recognition appeared only in the corpus-loaded run - One 2026-09-01 episode: a repository-free synthesis re-derived retained notes and reproposed rejected framings while the corpus-loaded session recognized them; an uncontrolled bundle, recorded as a starting point for better contrasts
- In-context learning presupposes context engineering - In-context learning only works when the right knowledge reaches the context window — the selection machinery that ensures this is itself learned and refined over deployment
- Information value is observer-relative - Information value is observer-relative: prior knowledge, tools, compute, and goals determine extractable structure, grounding use-shaped reshaping and discovery.
- Instruction specificity should match loading frequency - The loading hierarchy (CLAUDE.md → skill descriptions → skill bodies → task docs) should match instruction specificity to loading frequency — always-loaded context competes for attention every session
- KB goals in always-loaded context guide inclusion decisions - Without explicit goals in the always-loaded control-plane file, agents cannot reject well-written but off-scope material — a universal quality guide provides writing criteria but not domain scope
- Keep Lineage And Compiled Views From Drifting - Generated cues, prompt files, indexes, and assistant-specific views need lineage and authority rules so they do not drift into independent behavior-shaping force
- Knowledge storage does not imply contextual activation - Separates knowledge that exists, knowledge loaded into context (read-back), and knowledge that actually changes behavior (activation); explains why retrieval and long context do not guarantee activation
- Knowledge-access architecture must be evaluated end to end, not by retrieval alone - Explains why retrieval measures and storage-substrate labels cannot proxy for task-relative quality across discovery, loading, transformation, activation, and upkeep
- Legal drafting solves the same problem as context engineering - Legal drafting parallels context engineering because both write ambiguous natural-language specifications for judgment-based interpreters, but law develops constraining more than codification
- Link-following and search impose different metadata requirements - Compares contextual local steps with long-range search and explains why these recurring navigation modes impose different metadata requirements on an agent knowledge base
- Links encode conditional possibilities, not obligations - Links encode conditional possibilities, not obligations — every label must name a specific reader-need (the condition under which following pays off); content required for all reachable readers should be inlined, not linked
- LLM context is composed without scoping - Flat context concatenation lacks local scope and produces name collision, contamination, and spooky action at a distance; code-built sub-agent contexts must impose boundaries
- LLM contexts interpret instructions and content through the same token medium - LLMs interpret instructions and content through one token medium, enabling natural-language artifacts to alter behavior without translation while requiring architecture to enforce role, scope, and authority boundaries
- LLM recompute cost shifts the store-vs-recompute balance - For model-facing derived values, costly model-side recomputation shifts cache economics toward checked materialization, but persistence pays only when its total expected cost beats the alternatives and the copy substitutes for work
- LLM-mediated schedulers are a degraded variant of the clean model - When the agent scheduler lives inside an LLM conversation it becomes bounded and degrades; three recovery strategies — compaction, externalisation, factoring into code — restore the clean separation to increasing degrees
- Load-bearing vocabulary collisions should be prevented or visibly scoped at write time - Unqualified technical senses have no reliable namespace in natural-language content; schema slots, rare compounds, and linked clause frames scope them at write time; audits and remediation recover when prevention fails
- Local materialization should outperform distant natural-language declarations - Predicts that, for distant or non-obvious uses of a natural-language declaration, generated local materialization will outperform declaration-only presentation without creating a second maintenance authority
- Make Authority Explicit - Memory architecture must state who can read, write, promote, activate, enforce, revise, and retire memory across risk levels
- Memory design adds operational axes to artifact analysis - Memory design needs operational policy axes (capture, derivation, activation, authority assignment, lifecycle, evaluation) on top of substrate, form, lineage, and behavioral authority
- Memory-backed personalization can look like model improvement - Distinguishes user-specific gains supplied by retained intent from gains in the model that interprets the assembled context.
- Minimum viable vocabulary is the naming set that most reduces extraction cost for a bounded observer - Defines minimum viable vocabulary as the names that most reduce a bounded observer's extraction cost, connecting conceptual thresholds to an information-theoretic optimization
- Model-resolved indirection adds interpretation work to LLM execution - A reference adds model-side interpretation only when the model must resolve it; upstream literalization is worthwhile when binding, token, authority, and regeneration costs favor it
- Natural-language project state may specialize weight-resident search heuristics - The natural-language part of project state may specialize general search heuristics already represented in an LLM's weights by supplying current intent, theory, branch history, and constraints
- Naur's compiler case tests one historically bounded documentation-and-consumption system - Naur's compiler transfer failure rules out more documentation of the same kind, but tested one historically bounded package and consumption process rather than every possible rationale, indexing, retrieval, and activation system
- Naur's human-only conclusion needs more than the absence of explicit criteria - Naur's human-only conclusion needs a further premise connecting unformulated judgment to computational inability; this reading preserves his functional tests without claiming that learned criteria are inexpressible
- Open-domain memory retention needs a declared output spec - Explains why an input stream alone can't answer 'what to store' in open-domain memory design; a declared output spec supplies the missing inclusion criterion.
- Periodic KB hygiene should be externally triggered, not embedded in routing - Routing instructions serve the current task; periodic hygiene is triggered externally (user, heartbeat, CI), so embedding it in always-loaded routing blurs two responsibilities and adds session noise
- Pointer design tradeoffs in progressive disclosure - Compares fixed, query-time, and crafted retrieval pointers across specificity, cost, availability, accuracy, and authoring dependence
- Preserve Evidence Without Making History The Next Context - Trace retention should preserve evidence for audit and extraction without making raw history the agent's default context
- Promote Only When Future Value Exceeds Maintenance Cost - Candidate memory should become durable only when future retrieval or activation value exceeds review and maintenance cost
- Promotion selects for unreliable activation, and the regress ends only at an external trigger - Recasts promotion from 'the consumer lacks this' to 'the consumer will not apply this unprompted', and requires delivery to have a root firing event independent of that prior activation
- Raw accumulation does not create usable memory - Accumulation preserves material, but usable agent memory requires ingress work that adds handles, scope, relationships, provenance, trust signals, and lifecycle pressure.
- Retaining episode evidence keeps a distilled rule open to re-examination - Keeping relevant episode evidence and its relation to a distilled rule preserves a route for re-examining that rule; reconstruction, comparative value, and correct generalization still require testing
- Retire, Redact, Supersede, And Relax Memory - Memory systems need lifecycle operations for redaction, decay, supersession, retirement, relaxation, and temporal validity
- Rule-based context selection needs a pre-existing signal - A rule-based selector can target one case only when a rule-ready signal already distinguishes it; otherwise the system must wait, load broadly, or infer relevance from task and candidate content
- Semantic work can be relocated but not eliminated - A meaning-dependent judgment is never removed, only placed — moved upstream where its inputs exist (amortized) or off a bottlenecked context (offloaded); 'free at use-time' always means paid earlier
- Serve Multiple Consumers, Not One Retrieval Interface - Memory systems need multiple surfaces because acting, scheduling, review, learning, governance, and active work consume memory differently
- Session history should not be the default next context - Storing execution history and loading it into the next agent call are separate decisions; chat and framework-owned tool loops conflate them by making session history the default next context
- Seven documentation cases left routing and synthesis - A seven-artifact Commonplace sweep found that direct source access removed exact-fact prose while discovery maps and cross-component boundaries survived; it does not establish a universal documentation ratio
- Short composable notes maximize combinatorial discovery - The library's purpose is to produce notes that can be co-loaded for combinatorial discovery — short atomic notes are a consequence of this goal; longer synthesized artifacts belong in workshops or derived instructions
- Skill discovery re-fires in every sub-agent context, not just the top-level invocation - Skill discovery is per-context and autonomous — every installed skill is re-matched in each sub-agent context, even ones a parent narrowed, so a delegating skill's own discoverability is a leak vector
- Soft degradation can bind before the hard cap even when required evidence fits - For quality-sensitive agent work whose required evidence fits within the provider window, volume, complexity, and interference can silently constrain usable context before the hard cap
- Soft-bound traditions as sources for context engineering strategies - Survey of twelve soft-bound traditions as candidate sources for context engineering strategies, with a three-tier assessment of what transfers, what's plausible, and what's blocked
- Specification-level separation recovers scoping before it recovers error correction - OpenProse-like DSLs expose control flow and discretion boundaries while leaving scheduling and validation on the LLM substrate, creating an intermediate regime between flat prompting and symbolic scheduling
- Subtasks that need different tools force loop exposure in agent frameworks - When decomposition creates child tasks with different tool surfaces, the parent must construct fresh calls for each child, so a framework-owned loop is no longer the right control surface
- System-definition artifacts are crystallized reasoning under context scarcity - Separates heuristic rules that substitute for unavailable read-time reasoning from authority-bearing constraints and symbolic codification, which remain useful even with abundant context
- Tag maintenance and derived browsing - Proposal: test temporary topic groupings and reviewed tag-maintenance suggestions before changing Commonplace’s canonical tag structure.
- The adaptation survey corroborates memory requirements but misses artifact governance - The agentic-adaptation survey supports the memory requirements map by treating memory and skills as adaptive tools, but it needs substrate, form, lineage, and authority governance to become design guidance
- The chat-history model trades context efficiency for implementation simplicity - Chat history persists because appending messages preserves information and avoids interface design, but that convenience trades away selective loading under bounded context
- The four-field record exposes an efficiency, security, and sovereignty risk triad - The four artifact-analysis fields exist to surface three architectural review concerns over retained behavior — efficiency, security, and sovereignty — with sovereignty (owner control to inspect, regenerate, delete, roll back) as the new axis
- The practical scheduler is the host language, not a reified select - The simplest practical orchestration library demotes the tool loop to a returning, per-call-parameterized function and lets ordinary host-language code play select and K — reifying K only when the run must outlive its process or outgrow its memory
- Topology, isolation, and verification form a causal chain for reliable agent scaling - Topology, isolation, and verification may form a strict dependency chain rather than independent design choices — tested against the simpler account that good decomposition implies the other two
- Trace-extracted memory earns authority per operation, not at capture - Trace memories begin as records; verification, abstraction, and consultation earn authority under progressively harder oracles, while unverified stores accumulate guesses presented as knowledge
- Types give agents structural hints before opening documents - Types and descriptions let agents make routing decisions without loading full documents — the type says what operations a document affords, the description filters among instances of that type
- Under sub-agent decomposition, feasibility is the heaviest fork's net load - Shows why decomposition changes feasibility from total operation cost to the largest residual load left on any fork after work is shifted to siblings or the parent
- Use Trace Extraction As Meta-Learning - Trace extraction is an after-the-fact learning path that must respect signal quality, review, and readable-artifact versus distributed-parametric learning boundaries
- Writing styles are strategies for managing underspecification - Maps descriptive, prescriptive, prohibitive, explanatory, and conditional context-file styles to distinct ways of narrowing agent interpretation, each trading constraint against generality