Context engineering
Type: types/tag-readme.md
This tag gathers work on getting the right knowledge into a bounded LLM context at the right time: routing and retrieval, loading and prompt assembly, scoping, context budgets and degradation, scheduling work across context windows, and deciding what to store so it can be loaded later. The defining note is context engineering, which names routing, loading, scoping, and maintenance as the operational core. Members span kb/notes/, reference proposals for write-context assembly, and system analyses in kb/agentic-systems/. Nearby but different: agent-memory covers what an agent retains across sessions and under what authority; a note belongs here when its question is how knowledge reaches a bounded call, whether or not it was retained as memory. This head is selective; use a scoped tag search for full membership.
Core Claims
- Context engineering - the definition: the discipline of designing systems around bounded-context calls, with routing, loading, scoping, and maintenance as its core
- Context efficiency is the central design concern in agent systems - why context is the scarce resource, and why per-window degradation binds before token cost
- Design for the first-time human, except on access cost - explains why human-facing materializations and agent-facing query paths can share a source of truth while following different access modes
- semantic sub-goals that exceed one context window become scheduling problems - explains when context limits force orchestration instead of a single larger prompt
- stateful tools recover control by becoming hidden schedulers - shows how runtime state can relocate context control behind the tool boundary
- A context-operation interface bounds the projections its policy can realize - separates structural projection reach from policy success within one interface
- Addressability grain, not compression ratio, sets a matched selective-read floor - isolates the retrieval-floor condition for a known question that maps to one discriminating unit on each path, while leaving fan-out, sufficiency, reliability, and net value separate
- Opposed recompute factors do not decide documentation segmentation - why crossed savings and recurrence rankings do not order cache value, and why cache value alone does not decide whether a second content layer pays
- A derived copy of recomputable truth must be checked or absent - names when a recomputable value is safe to inline for context economy: only when a validator can re-derive and check it, otherwise it must stay a live read
Methodology Activation
- Capable agents need methodology selection - relevant but incompatible approaches force choosing a governing methodology, not just supplying knowledge
- Weight-resident methodologies compress behavior in context - a compact cue can activate a methodology already in the weights, at the cost of exact specification
Write-Context Assembly
- Deterministic write-context assembly - proposal: code assembles a target's fixed authoring context with closed input roles
- ADR 092: Write briefs are optional sidecars named by a validated pointer - decision: keep an artifact-specific commission beside its document and deliver it to writers
Related Tags
- Agent memory - the retention side: what agents keep across sessions and how it is activated; most memory-requirements notes carry both tags
- Computational model - the bounded-call substrate context engineering operates on, including framework-owned loops and when scheduling must become explicit
- Deploy-time learning - the post-release change that context machinery helps absorb without retraining
Other tagged notes
- A citation cannot assert more fidelity than its capture preserved - Capture is layered (verbatim / paraphrase / second-hand) by forced constraints; a citation's fidelity is bounded by which layer holds the passage, and no notation can raise it — only re-capture
- A compact, refreshable whole-picture narrative can replace infeasible fragment reconciliation - Holistic rewrite shifts reconciliation from each consumer to the author, but only when the whole-picture narrative can fit within effective context and be refreshed before the narrative goes stale
- A knowledge base should support fluid resolution-switching - Defines resolution-switching as movement among KB views with different scope and detail, then inventories the mechanisms and limits of that qualitative criterion
- A linked note's durable payload is what its consumption path cannot reliably supply - Retain the recognition anchor and rationale the intended consumption path cannot reliably supply — an enforced path can carry the anchor itself; reconstructable framework recap factors into the linked artifact, tested by downstream effects
- A retained instruction preserves what testing selected - Explains why an instruction generated from model weights can still add KB value: testing selects a procedure under a criterion and retention makes that choice reusable.
- A specific intent may out-yield local rationales, but contingent facts stay separate - Conjectures that an unrecoverable governing intent yields more local rationale per token than rationale snippets, while contingent design facts need their own record
- Academic Research Skills - Academic Research Skills as a prompt-defined Claude Code research pipeline with narrow executable checks, host-dependent orchestration, protocol-only resume, and conflicting terminal gate rules
- Activate Behavior-Changing Memory Before The Mistake - Behavior-changing memory must activate before relevant actions rather than waiting for explicit retrospective search
- Active work state is not retrospective memory or chat history - Active work state needs current pointers, evidence gates, and closure; treating it as retrospective memory or chat history preserves the wrong state
- Adaptation signals choose pressure; artifact analysis chooses the retained surface - Maps agentic-adaptation signals onto artifact-analysis axes so KB learning records which retained surface changes, what authority it gains, and how to review it
- Agent memory needs discoverable, loadable, composable, trusted knowledge under bounded context - Distinguishes four use-time requirements for remembered knowledge—discoverability, loadability, composability, and calibrated trust—from system-level activation.
- Agent Memory Requirements - Navigation hub for concrete agent-memory requirements extracted from the memory-system design synthesis
- AI Agents in Depth - Whole-book comparison of AI Agents in Depth with Commonplace, separating broad architectural convergence from differences in memory admission, epistemic warrant, governance, and orchestration
- An insufficient summary precedes the source rather than replacing it - When a summary cannot license a reliability-compliant stop, the authoritative fallback remains in the path; only fallback work the summary removes can offset its own cost
- Artifact function as a routing field - Proposal: whether an artifact_function declaration should expose a document's intended whole-artifact job for writing and review routing without asserting atomicity
- beads_rust - beads_rust as a local active-work and coordination substrate: transactional CLI claims and workflow gates, explicit external execution and Git boundaries, and weaker parity across MCP, inherited-context, and shipped instruction paths
- Borrowing can operate through retained artifacts or weight activation - Established external methodologies can become operative either by being explicitly retained in the system or by activating a model's pretrained representation; the two routes trade context economy against inspectability and revisability
- Bottom-up structure inference needs capture at the decision surface, not the state - Bottom-up inference of entities and relations from traces needs decision-shaped capture at the decision surface: the 'why' is cheap to record there and hard-to-impossible to recover from state later
- Brainstorming: how to test whether pairwise comparison can harden soft oracles - Staged test plan for whether pairwise comparison improves soft-oracle properties (discrimination, stability, calibration) in LLM evaluation loops
- Cheap generation breaks text volume as an effort signal - When text is cheap to expand but costly to verify, length stops evidencing author effort and can instead warn that the reviewer inherits unperformed checking
- Checked inline blocks for shared instruction text - Proposal: reuse natural-language authoring mechanics in specialized writer prompts through literal inlining backed by deterministic source-to-copy checks
- Context contamination operates below an agent's compliance reasoning - A controlled test found fine-grained stance drift despite explicit detection and refusal; exclusion guarantees non-exposure, while instruction-level mitigation remains an empirical question
- Create Memory Directly - Direct memory creation preserves live understanding by writing useful artifacts before later trace extraction loses structure
- Cross-task transition policy remains scheduling behind a tool interface - Classifies code by authority over interceptable transitions among independently steerable goals, separating scheduler role from its tool-shaped interface and audience-relative concealment
- Design rationale must preserve decision premises its interpreter cannot regenerate - Retention test for source-checkout design rationale: keep current decision premises not faithfully recoverable from implementation, git, and general knowledge; treat recoverable, role-free explanation as a cache
- Designing a Memory System for LLM-Based Agents - Derives agent-memory design pressures and links to a requirements inventory for agents designing or evaluating memory systems
- Evaluate Memory By Effects, Not By Existence - Memory should be evaluated by downstream effects on tasks, artifacts, answers, behavior, context efficiency, and lineage alignment
- History has one chance to become checkable - An artifact's production history is convertible to later-checkable form only at production time, via records/attestation or re-derivability; after that a bounded reviewer sees only carried state
- Import External Knowledge Into Internal Form - Agent memory systems need import paths when authoritative project knowledge already exists outside the memory substrate
- In one episode, recognition appeared only in the corpus-loaded run - One 2026-09-01 episode: a repository-free synthesis re-derived retained notes and reproposed rejected framings while the corpus-loaded session recognized them; an uncontrolled bundle, recorded as a starting point for better contrasts
- In-context learning presupposes context engineering - In-context learning only works when the right knowledge reaches the context window — the selection machinery that ensures this is itself learned and refined over deployment
- Keep Lineage And Compiled Views From Drifting - Generated cues, prompt files, indexes, and assistant-specific views need lineage and authority rules so they do not drift into independent behavior-shaping force
- Knowledge-access architecture must be evaluated end to end, not by retrieval alone - Explains why retrieval measures and storage-substrate labels cannot proxy for task-relative quality across discovery, loading, transformation, activation, and upkeep
- LLM recompute cost shifts the store-vs-recompute balance - For model-facing derived values, costly model-side recomputation shifts cache economics toward checked materialization, but persistence pays only when its total expected cost beats the alternatives and the copy substitutes for work
- Make Authority Explicit - Memory architecture must state who can read, write, promote, activate, enforce, revise, and retire memory across risk levels
- Memory design adds operational axes to artifact analysis - Memory design needs operational policy axes (capture, derivation, activation, authority assignment, lifecycle, evaluation) on top of substrate, form, lineage, and behavioral authority
- Mixed epistemic status must be preserved below the document level - A document can combine observations, deductions, and plausible explanations; KB writing and review must retain which claims and transitions have which warrant.
- Natural-language project state may specialize weight-resident search heuristics - The natural-language part of project state may specialize general search heuristics already represented in an LLM's weights by supplying current intent, theory, branch history, and constraints
- Naur's compiler case tests one historically bounded documentation-and-consumption system - Naur's compiler transfer failure rules out more documentation of the same kind, but tested one historically bounded package and consumption process rather than every possible rationale, indexing, retrieval, and activation system
- Naur's human-only conclusion needs more than the absence of explicit criteria - Naur's human-only conclusion needs a further premise connecting unformulated judgment to computational inability; this reading preserves his functional tests without claiming that learned criteria are inexpressible
- Open-domain memory retention needs a declared output spec - Explains why an input stream alone can't answer 'what to store' in open-domain memory design; a declared output spec supplies the missing inclusion criterion.
- Preserve Evidence Without Making History The Next Context - Trace retention should preserve evidence for audit and extraction without making raw history the agent's default context
- Promote Only When Future Value Exceeds Maintenance Cost - Candidate memory should become durable only when future retrieval or activation value exceeds review and maintenance cost
- Promotion selects for unreliable activation, and the regress ends only at an external trigger - Recasts promotion from 'the consumer lacks this' to 'the consumer will not apply this unprompted', and requires delivery to have a root firing event independent of that prior activation
- Raw accumulation does not create usable memory - Accumulation preserves material, but usable agent memory requires ingress work that adds handles, scope, relationships, provenance, trust signals, and lifecycle pressure.
- Retaining episode evidence keeps a distilled rule open to re-examination - Keeping relevant episode evidence and its relation to a distilled rule preserves a route for re-examining that rule; reconstruction, comparative value, and correct generalization still require testing
- Retire, Redact, Supersede, And Relax Memory - Memory systems need lifecycle operations for redaction, decay, supersession, retirement, relaxation, and temporal validity
- Rule-based context selection needs a pre-existing signal - A rule-based selector can target one case only when a rule-ready signal already distinguishes it; otherwise the system must wait, load broadly, or infer relevance from task and candidate content
- Semantic work can be relocated but not eliminated - A meaning-dependent judgment is never removed, only placed — moved upstream where its inputs exist (amortized) or off a bottlenecked context (offloaded); 'free at use-time' always means paid earlier
- Serve Multiple Consumers, Not One Retrieval Interface - Memory systems need multiple surfaces because acting, scheduling, review, learning, governance, and active work consume memory differently
- Soft degradation can bind before the hard cap even when required evidence fits - For quality-sensitive agent work whose required evidence fits within the provider window, volume, complexity, and interference can silently constrain usable context before the hard cap
- Soft-bound traditions as sources for context engineering strategies - Survey of twelve soft-bound traditions as candidate sources for context engineering strategies, with a three-tier assessment of what transfers, what's plausible, and what's blocked
- Subtasks that need different tools force loop exposure in agent frameworks - When decomposition creates child tasks with different tool surfaces, the parent must construct fresh calls for each child, so a framework-owned loop is no longer the right control surface
- Tag maintenance and derived browsing - Proposal: test temporary topic groupings and reviewed tag-maintenance suggestions before changing Commonplace’s canonical tag structure.
- The adaptation survey corroborates memory requirements but misses artifact governance - The agentic-adaptation survey supports the memory requirements map by treating memory and skills as adaptive tools, but it needs substrate, form, lineage, and authority governance to become design guidance
- The practical scheduler is the host language, not a reified select - The simplest practical orchestration library demotes the tool loop to a returning, per-call-parameterized function and lets ordinary host-language code play select and K — reifying K only when the run must outlive its process or outgrow its memory
- Trace-extracted memory earns authority per operation, not at capture - Trace memories begin as records; verification, abstraction, and consultation earn authority under progressively harder oracles, while unverified stores accumulate guesses presented as knowledge
- Use Trace Extraction As Meta-Learning - Trace extraction is an after-the-fact learning path that must respect signal quality, review, and readable-artifact versus distributed-parametric learning boundaries
- Warranted reader update is the objective of substantive writing - Defines epistemic interestingness as a relevant, warranted change relative to an intended reader's prior, making contribution selection—not accumulated inputs—the purpose of multistage writing.