Notes Directory
Type: kb/types/generated-index.md
← Parent
Subdirectories
Files
- "Agent" is a useful technical convention, not a definition (note) - A lightweight technical convention — an agent is a tool loop (prompt, capability surface, stop condition) — sidestepping the definitional debate in favor of a unit that organizes code
- A bare writing prompt does not determine its intended contribution (note) - Separates the contribution a bare writing prompt leaves underdetermined from empirical claims about how experts and LLMs supply the missing purpose.
- A benchmark that holds the client fixed exports the least-warrantable decisions by design (note) - A fixed-client benchmark measures worker capability; it leaves broader closure untested when the client supplies internal production decisions, while ordinary user requirements and acceptance may remain external
- A better-factory claim compares operative states under an antecedent assessment relation (note) - The improvement claim's relata are predecessor and operative-successor states and its relation is declared before the development it judges; evaluator location is a separate declaration from the learner boundary
- A borrowed pattern transfers only as far as source and target share a mechanism (note) - A borrowed pattern carries transferred warrant only over the layer where source and target share the mechanism it depends on; where the link is analogy or the sharing doesn't reach, it must earn adoption by target-side evidence
- A capable agent needs methodology selection, not just relevant knowledge (note) - A capable agent may know many individually relevant but mutually incompatible approaches, so task control requires selecting a governing methodology rather than merely supplying relevant knowledge
- A checked outcome licenses retaining an episode, not abstracting its explanation (note) - One result-only check can warrant retaining an episode as evidence, but abstracting its explanation also needs evidence about a faithful producing process and an explicit scope boundary
- A citation cannot assert more fidelity than its capture preserved (note) - Capture is layered (verbatim / paraphrase / second-hand) by forced constraints; a citation's fidelity is bounded by which layer holds the passage, and no notation can raise it — only re-capture
- A claim's warrant does not determine its fit in a working theory (note) - Independent warrant and fit in a working theory answer different questions: a warranted claim may fit poorly, while apparent fit may be produced by an unwarranted or already-assumed claim
- A compact, refreshable whole-picture narrative can replace infeasible fragment reconciliation (note) - Holistic rewrite shifts reconciliation from each consumer to the author, but only when the whole-picture narrative can fit within effective context and be refreshed before the narrative goes stale
- A consumption channel delivers force without the history that earned it (note) - A consumption path can promote content into a higher-force role without checking whether an authorization covers that content, version, and use
- A context-operation interface bounds the projections its policy can realize (note) - Explains why improving context selection within a fixed operation interface cannot establish that the interface admits every useful active-context projection.
- A derived copy of recomputable truth must be checked or absent (note) - When an artifact carries a copy of information recomputable from a ground-truth source, the copy must be machine-checked against that source or not exist — hand-maintained-and-trusted is forbidden
- A failure explanation becomes search control only when it changes a later branch decision (note) - An explanation of a failed branch becomes operative search control only when its retention changes a later choice about scope, priority, probing, continuation, or abandonment
- A fixed-model house must retain missing procedures for theory use (note) - With models pinned, newly acquired theory-use procedures must persist outside their parameters; existing general machinery may already supply them, while code can make specified steps cheaper and more reliable
- A framework rule with a boundary-preserving rival is not an inherited constraint (note) - A rival design that preserves a framework's boundary invariants while dropping a rule demotes the rule to a design choice; finding no rival certifies nothing — the test cuts one way only
- A functioning knowledge base needs a workshop layer, not just a library (note) - Explains workshop as the temporal counterpart to a permanent knowledge library: in-flight state, dependencies, expiry, and promotion bridges, with tasks as an early prototype
- A goal-holding interpreter fails soft, and its workarounds tax a bounded budget (note) - A procedure compiles its goal away, so a blocked step fails loud and hard; an interpreter holds the goal and re-routes, so failures are absorbed as a per-encounter tax on bounded capacity — silent, accumulating, and softly saturating
- A hand-crafted bootstrap fits the Bitter Lesson only if learning can outgrow it (note) - A hand-crafted starting state fits the Bitter Lesson only if scalable learning displaces the task- and family-specific production knowledge it supplies as claimed reach widens
- A knowledge base should support fluid resolution-switching (note) - Defines resolution-switching as movement among KB views with different scope and detail, then inventories the mechanisms and limits of that qualitative criterion
- A linked note discharges its own grounding, so a citing note owes representation, not re-grounding (note) - A cited source imposes a grounding obligation; a claim-titled note that already passed its own grounding review imposes only a representation obligation — with the preconditions that keep the distinction and why it is not a paraphrase ledger
- A linked note's durable payload is what its consumption path cannot reliably supply (note) - Retain the recognition anchor and rationale the intended consumption path cannot reliably supply — an enforced path can carry the anchor itself; reconstructable framework recap factors into the linked artifact, tested by downstream effects
- A method's ceiling bounds the method, not the transfer it already made (note) - Separates envelope expansion, where a responsibility leaves the residual human work, from performance gains inside a fixed envelope, so a bounded method reaching its ceiling does not retract the transfer it already made
- A methodology governs its own extension only as far as it settles the meta-decisions it raises (note) - A retained methodology governs the consequential extension decisions it supplies or imports; actor competence can carry the process further without making those choices settled by the method
- A note is an atomic step relative to the check that reads it (note) - Two independent bounds on a note: one claim sized to the reader's bounded context, and one checkable inference sized to the checker's single pass — for the grounding check the unit is the unquoted source
- A proposal-selection improvement loop requires search, evaluation, and operative retention (note) - A proposal-selection improvement loop — candidates generated, evaluated with possible non-adoption, and accepted changes made operative — requires search, reject-capable evaluation, and operative retention
- A proximate target is checked for achievement, not for warrant (note) - Between an improvement objective and its oracles sits a target level — a property pursued because it is held to serve the objective — whose linking claim no check in the loop evaluates
- A repeatable operative path keeps a redesign class open to revision (structured-claim) - Operationalizes repeatable operative revision for a named redesign class as a causal path through representation, evidence-bearing determination, admission, installation, dependence, and continuity
- A retained instruction preserves what testing selected (note) - Explains why an instruction generated from model weights can still add KB value: testing selects a procedure under a criterion and retention makes that choice reusable.
- A retained-theory intervention isolates one surface, not the whole program theory (note) - An intervention on retained theory estimates that surface's causal contribution under matched conditions; influence, explanatory guidance, acquisition, and whole-system theory possession remain different claims
- A retrieval miss is a local reflective-path failure (note) - A missed relevant artifact leaves its represented aspect inert for the affected task and discovery route, while other loading paths and reflective aspects can remain causally connected
- A search controller is tested by what it brings to stronger evaluation (note) - A search controller should be evaluated by the branches and probes it routes into stronger evaluation, not by treating every provisional judgment as an acceptance claim
- A software factory is family-scoped lifecycle production machinery (note) - Reconstructs Greenfield's versioned software-factory ontology: declared product family, schema, packaged assets, configured environment, two development processes, and lifecycle work products
- A specific intent may out-yield local rationales, but contingent facts stay separate (note) - Conjectures that an unrecoverable governing intent yields more local rationale per token than rationale snippets, while contingent design facts need their own record
- A theory's prototype standing is its revision cost: external binding plus lost investment (note) - A theory's prototype standing is its expected revision cost — external binding plus the investment a revision discards — so natural-language versus symbolic form determines neither component and acceptance status is a separate axis
- A universal knowledge framework demotes content taxonomies to defaults (note) - Universal frameworks should replace closed content taxonomies with complete local contracts and guarded creation-time defaults; what stays fixed is stipulated or enforced, not certified universal
- Abstract an experience into a lesson only when you can state where the lesson stops (note) - Abstract an episode into a lesson only when you can state its boundary, else preserve the instance; an over-generalized lesson is one that drops the condition clause
- Access burden and transformation burden are distinct query dimensions (note) - Separates the system-relative cost of finding required inputs from producing an answer, so query systems can diagnose which work remains as retrieval and reasoning interact
- Accumulation counts dependence through the retained result, not through the evidence it caused (note) - Cumulativity counts dependence through the retained result only; counting the evidence channel that result caused would make it coextensive with operativity
- Active work state is not retrospective memory or chat history (note) - Active work state needs current pointers, evidence gates, and closure; treating it as retrospective memory or chat history preserves the wrong state
- Ad hoc explanation can be rational when error is cheap and local (note) - Explains why a disposable local guess can rationally select the next probe when error is cheap and contained, while retained explanations need reach checks
- Ad hoc prompts extend the system without schema changes (note) - Any system with an LLM agent layer can absorb new requirements through natural language prompts without changing the deterministic base
- Addressability grain, not compression ratio, sets a matched selective-read floor (note) - For recoverable content and a known one-unit question, a summary lowers the matched raw read-volume floor only when its path loads less answer-bearing material than the source path; whole-artifact compression alone does not decide that relation
- Agent memory (tag-readme) - Curated head for agent-memory notes — memory as crosscutting architecture, requirements, activation, lifecycle, and evaluation
- Agent memory is a crosscutting concern, not a separable niche (note) - Memory decomposes into storage (solved), retrieval/activation (context engineering), and learning (learning theory) — treating it as a standalone category hides that the hard problems are at the intersections
- Agent memory needs discoverable, loadable, composable, trusted knowledge under bounded context (note) - Distinguishes four use-time requirements for remembered knowledge—discoverability, loadability, composability, and calibrated trust—from system-level activation.
- Agent orchestration needs a privilege quarantine, not just a permission scope (note) - When one agent in an orchestration reads untrusted content, the defense is a role-level privilege quarantine — barring that agent from high-privilege actions entirely — not finer per-call tool scoping
- Agent orchestration needs coordination guarantees, not just coordination channels (note) - Coordination channels say how bounded contexts interact, but the missing discriminator is which guarantee prevents contamination, inconsistency, amplification, or liability diffusion across the composed system
- Agent orchestration occupies a multi-dimensional design space (note) - Agent orchestration is not ordered along a single ladder — scheduler placement, persistence, coordination form, coordination guarantees, and return artifacts vary independently across architectures
- Agent statelessness makes routing architectural, not learned (note) - Because each session starts without learned navigation intuition, skills, type templates, routing tables, names, and activation triggers remain permanent architecture rather than temporary scaffolding
- Agent statelessness means the context engine should inject context automatically (structured-claim) - Since agents can't carry vocabulary or decisions between reads, the context engine should auto-inject referenced context — definitions once per session, ADRs when relevant. The trigger mechanism is open; the need follows from statelessness
- Agent-runtime analysis should separate scheduling, context assembly, and external state (note) - For runtimes composed of bounded model calls, separating control progression, per-call context, and external state or action services localizes failures even when one implementation owns all three
- Agentic systems interpret underspecified instructions (note) - Separates semantic underspecification from execution indeterminism: natural-language specs admit multiple valid projections, while constraining commits one projection to precise code.
- Agents navigate by deciding what to read next (note) - Models agent navigation as repeated follow/skip judgment under bounded context: cue diagnosticity must repay its own context cost, so longer pointer context is not automatically better
- AGENTS.md should be organized as a control plane (note) - Theory for deciding what belongs in AGENTS.md using loading frequency and failure cost, with layers, exclusion rules, and migration paths
- Alexander's patterns connect to knowledge system design at multiple levels (note) - Maps Alexander's Context/Problem/Forces/Solution pattern to typed document contracts and his generative process to incremental codification, while marking the looser 'centers' analogy
- Always-loaded context mechanisms in agent harnesses (note) - Survey of always-loaded context mechanisms across agent harnesses — system prompt files, capability descriptions, memory, and configuration injection — cataloguing what each carries, how write policies differ, and where the gaps are
- An accepted edit verifies the change, not the rule (note) - Human acceptance of an edit is a strong oracle for 'this change was wanted here' but a weak oracle for 'this generalizes' — mining rules from accepted edits inherits instance-level verification while the generalization step stays oracle-poor
- An action model matters only through its consumption path (note) - Agentic action can be direct or model-mediated; a retained action model matters only when its consumption path affects intervention selection
- An adversarial human-agent loop can reconstruct the writing-is-thinking filter (note) - The writing-is-thinking filter is the loop's, not the pen's — an adversarial human-agent loop can reconstruct what naive delegation loses, but only while the human stays the judge
- An agentic substrate becomes a software factory through family-specific production machinery (note) - Maps the bounded-call agentic substrate to Greenfield's software-factory ontology without calling every generic harness or generated program a factory
- An artifact must preserve the scope of each named system choice (note) - An artifact may inherit scope from context guaranteed to its consumers; for each named system choice it must preserve a proposition-relative reference rule or range plus the choice's role, not necessarily concrete identity or quantifier syntax
- An author should fix what the executor can't determine, not what it will (note) - A decision-specific rule for fixing what verified doctrine, task intent, and authorized evidence cannot safely determine while leaving bounded choices to execution
- An enforced tag-README combines a MOC pattern with checked membership (note) - A Commonplace tag-README can inherit Milo's contextual mapping pattern while validation checks only its declared membership relations, not editorial quality.
- An experiment identifies only the contrast it actually runs (note) - Why missing comparisons, bundle-to-component attribution, and adjacent unrun treatments all overstate causal conclusions beyond an experiment's observed contrast
- An insufficient summary precedes the source rather than replacing it (note) - When a summary cannot license a reliability-compliant stop, the authoritative fallback remains in the path; only fallback work the summary removes can offset its own cost
- An omitted improvement-loop function and a frozen one need different repairs (note) - Five proposal-selection systems expose frozen functions, while a direct-update contrast shows why absence of a gate is not omission; HyperAgents supplies a preliminary partial unfreezing
- An open-domain theory builder becomes a software house when new domains require production-machinery changes (note) - A persistent automated theory builder for external users becomes a software house when genuinely new domains require it to revise the software that performs theory production rather than only the theories produced
- Any barrier-delimited symbolic program with LLM calls is a batched select/call program (note) - Explains the LLM-call projection preserved when symbolic programs are converted to batched select/call form, and the additional semantics needed for effects and mutable harness configuration
- Architecture (tag-readme) - How Commonplace is structured and installed — repo layout, two-tree split, control-plane design, file-based storage
- Areas exist because useful operations require reading notes together (note) - Explains why orientation and comparative reading need bounded, sufficiently related note sets, while fixed sizes, membership rules, tags, and index layouts remain implementation choices
- Artifact classification separates content kind, lineage, and authority (note) - Use this note to classify retained KB artifacts without conflating content kind, production lineage, or path-relative behavioral authority with the collection's local writing contract.
- artifact-analysis (tag-readme) - Curated head for the artifact-analysis tag — the four-field vocabulary (substrate, form, lineage, authority) for classifying retained behavior-shaping artifacts, plus its applications
- Attempted recovery identifies informational gaps, not provenance or authority (note) - Recovery failure shows content is missing from the tested source; causal provenance and live authority require independent evidence, and only pairs with unique content on both sides are bidirectionally irrecoverable
- Automated synthesis is missing good oracles (note) - Generating synthesis candidates (cross-note connections, novel combinations) is easy — LLMs do it readily. The hard part is evaluating whether a candidate is genuine insight or noise.
- Automated tests for text (note) - Text artifacts can be tested with the same pyramid as software — deterministic checks, LLM rubrics, corpus compatibility — built from real failures not taxonomy
- Automating KB learning is an open problem (note) - The KB already learns through manual improvement; automating judgment-heavy mutations needs oracles for connections, groupings, and synthesis we cannot yet manufacture
- Axes of artifact analysis (note) - Artifact analysis records retained behavior-shaping artifacts by storage substrate, representational form, lineage, and behavioral authority so review evidence, invalidation, and rollback follow how artifacts actually act
- Backtracking keeps lightweight search control provisional (note) - Backtracking preserves the provisional status of a heuristic branch choice by restoring an earlier usable state and redirecting search after contrary evidence
- Borrowing can operate through retained artifacts or weight activation (note) - Established external methodologies can become operative either by being explicitly retained in the system or by activating a model's pretrained representation; the two routes trade context economy against inspectability and revisability
- Bottom-up structure inference needs capture at the decision surface, not the state (note) - Bottom-up inference of entities and relations from traces needs decision-shaped capture at the decision surface: the 'why' is cheap to record there and hard-to-impossible to recover from state later
- Bounded-context orchestration model (note) - Defines a conditional batched select/call form for closed-world LLM orchestration, including its state, barrier, feasibility, and comparison boundaries
- Brainstorming how to enrich web search (note) - Design exploration for enriching web search by reusing /connect's dual discovery and articulation testing on results, building a temporary research graph before bridging to KB
- Brainstorming: how explanatory-reach informs KB design (note) - Deutsch's reach, registered here as explanatory-reach, applied to KB notes — a maintenance risk signal, not a retrieval signal, because high-explanatory-reach revisions break downstream reasoning silently
- Brainstorming: how to test whether pairwise comparison can harden soft oracles (note) - Staged test plan for whether pairwise comparison improves soft-oracle properties (discrimination, stability, calibration) in LLM evaluation loops
- Brainstorming: maintainability oracles for agentic development (note) - Explores candidate signals, calibration experiments, authority levels, and workflow placements for evaluating maintainability in agent-generated code
- Broad software demands create pressure for agentic factory development (note) - Broad software demands make exhaustive predefinition of useful family-specific production machinery practically implausible, motivating agentic factory development without ruling out a fixed universal substrate in principle
- Candidacy evidence licenses escalation to assessment, not acceptance (note) - Separates candidacy evidence, which routes a hypothesis to costly assessment, from verdict evidence, which decides it; pricing and source-grounding cases provide two worked witnesses
- Canonical files may defer a shared schema while database authority remains a separate commitment (note) - Canonical files can defer a centralized schema while meanings remain unsettled; a database becomes canonical only when an explicit authority decision and operative write path commit resolutions the files no longer determine.
- Capability placement should follow autonomy readiness (note) - Three tiers — skills (autonomous-ready), instructions (reusable-but-steered), methodology notes (exploratory) — keep AGENTS.md free of capability inventories with a clear promotion path
- Causal and proof obligations are two formal routes to assessing explanatory-reach (note) - Causal and proof obligations demonstrate two ways formal symbolic systems can assess explanatory-reach inside a warranted model
- Changing requirements conflate genuine change with disambiguation failure (note) - Separates world change from late discovery that downstream work chose the wrong interpretation of an underspecified requirement; short iterations mainly limit propagation of the latter
- Cheap adoption and weak retirement accumulate cost (note) - Explains why locally cheap structural additions become routing, maintenance, and migration debt when retirement needs distributed evidence and lacks an equally operative path.
- Cheap generation breaks text volume as an effort signal (note) - When text is cheap to expand but costly to verify, length stops evidencing author effort and can instead warn that the reviewer inherits unperformed checking
- Choosing what to learn requires both validity and learning-value gates (note) - Separates two promotion checks for learning loops: whether a candidate is trustworthy enough to learn from, and whether learning it would improve the current system.
- Citing retained theory at the decision point is a mediation trace (note) - A decision record that cites the theory it followed supplies cheap, checkable evidence that the theory was consumed — necessary for a record-based mediation claim, but short of showing correct or load-bearing use
- Claim modality is the inference form of the refuter (note) - The three claim modes are refuter-defined images of deduction, induction, and comparative abduction; grounds the mode list's closure for empirical claims and gives vacuity and genre drift precise readings
- Claim notes should use Toulmin-derived sections for structured argument (structured-claim) - Three independent threads converged on Toulmin's argument structure — adopting Toulmin sections as base type
structured-claimseparates claim-titled notes (any note) from fully argued claims (the type) - Claim-routed reading may beat reading-first for synthesis notes (note) - Conjecture: writing a provisional claim first and reading only passages likely to overturn it may build a better-warranted synthesis note at lower context cost than reading everything first — motivated by Karnofsky, untested here.
- Claw learning loops must improve action capacity, not just retrieval (note) - A Claw learning loop must target contextual competence (execution, classification, planning, communication), not just retrieval accuracy — question-answering is one mode among many
- Code complements the weight–prompt pair with independently executed symbolic operations (note) - A model-mediated operation is instantiated by weights plus prompt; code complements that pair by defining operations whose consequences a symbolic runtime executes without reinterpreting the prompt
- Codification and relaxing navigate the bitter lesson boundary (note) - Since you can't identify which side of the bitter lesson boundary you're on until scale tests it, practical systems must codify and relax — with spec mining avoiding the vision-feature failure mode
- Codify-versus-LLM decision heuristics (note) - Specification, checking, permitted interpretations, and repeated-use cost inform which operations to codify; none alone makes code or model interpretation universally preferable
- Commitment, not derivation, creates new ground truth (note) - Derivation — claims recoverable from the source, nothing added — leaves the source as ground truth; what adds unentailed resolutions becomes ground truth at commit, repaired by supersession
- Compiling a coordination strategy preserves primitive authority but expands aggregate authority (note) - Compiling a coordination strategy preserves the primitive action alphabet but expands aggregate authority — the single-context envelope it escapes bounded both compute and effect volume
- Compounding is tested in later improvement, not by the accepting metric (note) - Compounding evidence must come from later improvement episodes through displaced productivity measures and causal traces, not from the metric that accepted the earlier change
- Computational model (tag-readme) - Tag README — PL concepts (scoping, homoiconicity, partial evaluation, typing) applied to LLM instructions, plus the scheduling architecture that follows from context scarcity
- Computationally directed self-improvement is a fixed-boundary reallocation ending in contraction (note) - The progress question for self-improving systems is not category membership but which decision-bearing functions humans still supply; the endpoint test is whether the boundary can be contracted to exclude them
- constraining (tag-readme) - Curated head for the constraining tag — narrowing the interpretation space of artifacts, from conventions to deterministic code; codification, relaxing, and the decision heuristics
- Constraining and extraction can trade generality for reliability, speed, or cost (note) - Constraining narrows interpretation and extraction produces focused use-shaped artifacts; both can trade generality for reliability, speed, or cost when task fit is good
- Constraining during deployment is continuous learning (note) - Continuous learning can happen outside of weights; constraining is one symbolic-artifact form where prompts, schemas, tools, and tests accumulate durable adaptive capacity during deployment
- Context contamination operates below an agent's compliance reasoning (note) - A controlled test found fine-grained stance drift despite explicit detection and refusal; exclusion guarantees non-exposure, while instruction-level mitigation remains an empirical question
- Context efficiency is the central design concern in agent systems (note) - Context is the single scarce resource in agent systems, and it is scarce for two distinct reasons — per-window degradation (feasibility) and aggregate token economics (cost) — of which feasibility is the binding one
- Context engineering (tag-readme) - Index for context-engineering notes about selecting, scoping, and maintaining task-relevant knowledge under bounded context
- Continual learning requires governing behaviour-changing writes, not just storing content (note) - For deployed systems, persistence is insufficient; continual learning must select, validate, authorize, and coordinate behaviour-changing updates across the representational forms a system can change
- Conversation vs prompt refinement in agent-to-agent coordination (note) - Conversation preserves the execution trace; prompt refinement compresses it into a clean handoff. The right choice depends on architecture and how much intermediate work should survive
- Convert still requires semantic description
- Cost-sensitive formalisms for fallible theory search (note) - Exploratory map of backtracking, learning, and complexity models that expose budgets relevant to fallible theory-guided search
- Criteria edits invalidate verdicts; process edits invalidate artifacts (note) - Editing quality criteria invalidates verdicts; editing production processes calls for artifact regeneration — verdict freshness includes the artifact and criteria but excludes its production process
- Cross-task transition policy remains scheduling behind a tool interface (note) - Classifies code by authority over interceptable transitions among independently steerable goals, separating scheduler role from its tool-shaped interface and audience-relative concealment
- Current-task fit alone does not warrant costly structural entrenchment (note) - Distinguishes reversible adoption from costly structural entrenchment and confines option reasoning to the timing of a commitment supported by an enduring constraint, scoped transfer warrant, or actual coordination value.
- Decomposition heuristics for bounded-context scheduling (note) - Working heuristics for symbolic scheduling over bounded LLM calls — separate selection from joint reasoning, choose representations not just subsets, save reusable intermediates in scheduler state
- Decorrelated reviewers still share the field's prior, so read their findings by the claim's stance (note) - Decorrelating reviewers removes the author's errors, not the field's; judges converge on consensus where a claim is original, so findings are read by stance and as reconnaissance: engage where load-bearing, deflect where not
- deploy-time-learning (tag-readme) - Curated head for the deploy-time-learning tag — the framework of system adaptation through durable, inspectable artifacts, plus learning fundamentals and feedback quality
- Derivation and inheritance give starting warrant; discriminating evidence or proof earns scope (note) - For reusable decompositions, derivation and inheritance supply conditional or transferred starting warrant, while evidence or proof earns only the scope it covers
- Descriptive link labels may supply the self-sufficiency a reconstruction gate would check (note) - Conjecture, partially tested: a claim-reconstruction gate is redundant on Commonplace notes; a label-ablation test attributes the self-sufficiency to the body-premise convention, not link labels. The thin pre-connect-draft case stays untested.
- Design for the first-time human, except on access cost (note) - Uses competent-newcomer ergonomics as a property-by-property default, then isolates access paths that charge a consumer for a selected slice or a whole artifact
- Design rationale must preserve decision premises its interpreter cannot regenerate (note) - Retention test for source-checkout design rationale: keep current decision premises not faithfully recoverable from implementation, git, and general knowledge; treat recoverable, role-free explanation as a cache
- Designing a Memory System for LLM-Based Agents (note) - Derives agent-memory design pressures and links to a requirements inventory for agents designing or evaluating memory systems
- Deterministic validation should be a script (note) - Half of /validate's checks are hard-oracle (enums, link resolution, frontmatter structure) and could run as a Python script in milliseconds instead of burning LLM tokens via the skill
- Diagnostic richness constrains outer-loop learning quality (note) - Outer-loop learning depends on inspectable failure evidence, not only on the oracle used to select winning candidates
- Directory placement is total, frontmatter classification is partial (note) - Canonical paths cover every file before validation and supply locality; opt-in types supply portability. Validation can encode similar policy on either surface, but native guarantees differ.
- Directory-scoped types are cheaper than global types (note) - Globally eligible types widen every collection's authoring choices; collection-local types keep specialized contracts scoped while path pointers load either kind on demand
- Discarding all experience-dependent state prevents cross-run accumulation (note) - Discarding an intermediate artifact loses that artifact's reuse path, not all learning; cross-run accumulation fails only when no experience-dependent state survives to affect later work
- Discarding software requires preserving the operational knowledge later work needs (note) - Regeneration must preserve or recover the commitments later operation depends on; reuse, reconstruction cost, and failure consequences matter more than code size or explanatory-reach alone
- Disconnected witnesses do not establish a full causal path through theory (note) - Theory use, outcome, theory revision, and later use establish theory-mediated learning only when their witnesses identify the joins of the same full causal path
- discovery (tag-readme) - Curated head for the discovery tag — positing a general concept and recognizing particulars as its instances; explanatory-reach as the value of what discovery produces
- Distinct residue classes require distinct functions in a self-improving architecture (note) - Different reasons for an untransferred decision identify different missing functions; a single process can supply several, and the current carrier split is not a permanent requirement
- Document system (tag-readme) - Index of notes about document types, writing conventions, validation, and structural quality — how notes are classified, structured, and checked
- Document types should be verifiable (note) - Document types should assert verifiable structural properties, not subject matter — with a base type + traits model inspired by gradual and structural typing
- Domain pricing routes an exception to idealization assessment but does not decide it (note) - Pricing signatures are defeasible, author-external evidence that a counterexample deserves idealization assessment; whether it refutes is settled by intended use, the omitted mechanism, consequence bounds, and explanatory dominance
- Edge ownership selects the key; choosing files or a database requires a workload comparison (note) - Distinguishes the complete edge key required by relation-owned mutable state from the workload-specific choice between edge files and a database.
- Elicitation requires maintained question-generation systems (note) - Four elicitation strategies ordered by user expertise required, composable into review architectures with maintenance loops that prevent ossification
- Enforcement without structured recovery is incomplete (note) - The enforcement gradient covers detection and blocking but has no recovery column — recovery strategies (corrective → fallback → escalation) are the missing layer, and oracle strength determines which are viable at each level
- Epiplexity by example: what entropy and complexity miss (note) - ELI5 explanation of epiplexity through encrypted messages, shuffled textbooks, CSPRNGs, and chess notation — contrasting surprise, shortest description, and observer-relative usable structure
- Error correction works with above-chance oracles and decorrelated checks (note) - Error correction for LLM output is viable whenever the oracle has discriminative power (TPR > FPR) and checks are decorrelated — amplification cost scales with 1/(TPR-FPR)² and independence of errors
- Error messages that teach are a constraining technique (note) - In agent systems the error channel is an instruction channel — making errors teach the fix is nearly free and eliminates the agent's need to diagnose, an orthogonal axis to enforcement strength
- Evaluation (tag-readme) - What works, what doesn't, what needs testing — empirical observations about KB operations and prompt design
- Evaluation automation is phase-gated by comprehension (note) - Optimization loops need diagnostic error analysis and demonstrated judge discrimination before automation can improve behavior rather than just score
- Exact implementation does not validate a requirement against its objective (note) - An artifact can exactly implement a requirement while the requirement remains a conjectured proxy for a declared objective; assess each named path separately, and attribute failure to the link without erasing local correctness
- Execution indeterminism is a property of the sampling process (note) - The same prompt can produce different outputs across runs due to token sampling — this is a property of the execution engine, theoretically eliminable but practically ubiquitous, and often confused with the deeper issue of underspecification
- Execution shaping determines directory placement (note) - Hunch that artifacts shaped as executable procedures belong in kb/instructions/ — the directory boundary is execution form, not compression or loading frequency
- Explicit retention provides direct targets for selective revision (note) - Explicit artifacts give a learner direct targets for inspecting and revising commitments; durability, writability, and effective addressability still depend on the boundary and available operations
- Factory construction is not evidence of production-knowledge acquisition (note) - Recursive software-factory construction is prior art, but the demonstrated constructors receive the family definitions, metamodels, mappings, and expertise that determine the produced factory
- Factory learning is experience-responsive retention that improves the factory (note) - Experience-responsive retention: production experience determines a retained change to reusable family machinery that later production depends on; factory-level learning is retention that improves the factory relative to a declared objective
- Factory-learning mechanisms should be compared on the same causal job (note) - Compares factory-learning mechanisms on their shared causal job — experience-responsive retention — while separating update mechanisms from the project-theory function needed for open-ended coherent modification
- Failure modes (tag-readme) - Index for failure-modes notes about characteristic ways knowledge can exist without changing agent behavior
- False-positive generation is filtered; false-positive acceptance becomes operative (note) - False-positive generation faces evaluation before retention, while false-positive acceptance becomes operative and can compound
- Feedback-trained memory management is oracle-dependent even when its operations are hand-designed (note) - Fixed and merely runtime-responsive memory rules need no training oracle; outcome-driven updates do, while noisy rankings weaken learning and misaligned ones teach the wrong ordering
- Final task success does not establish intended-path health (note) - When intended-path success and fallback success produce the same final task outcome, the outcome cannot establish whether the intended path and its supporting infrastructure were healthy
- First-principles analysis maps a design space before selecting within it (note) - Why deriving independent choice dimensions from boundary constraints exposes rival designs that inherited solution categories hide
- First-principles reasoning selects for explanatory-reach over adaptive fit (note) - First-principles reasoning selects explanations with explanatory-reach, accountable to observed fit, premise variation, and rival-practice tests
- Flat memory predicts specific cross-contamination failures that are empirically testable (note) - Flat memory predicts three cross-contamination failures — search pollution, identity scatter, insight trapping — testable via an observation protocol against real agent systems
- Foundations (tag-readme) - Core theory the rest of the KB builds on — contextual competence, bounded context, explanatory-reach, design methodology, composability
- Frontloading is partial evaluation, not divide-and-conquer (note) - The partial-evaluation framing for LLM frontloading is structurally precise, not metaphorical, because instructions and data share one token medium; without that shared medium, it would just be divide-and-conquer
- Frontloading spares execution context (note) - Pre-computing known instruction inputs and inserting their results spares execution-context budget inside a later LLM call
- Full-identity keys decouple a batch protocol from its packing axis (note) - A batched LLM-call protocol keyed by each unit's full composite identity, not position or a single axis, lets grouping strategy vary freely without protocol change
- Generality bought to avoid counterexamples is paid for in precision (note) - Widening a claim's vocabulary to survive counterexamples raises universality by spending precision, so content stays flat — and the unreadability that follows is the symptom, not the price of rigor
- Generate KB skills at build time, don't parameterise them (note) - Template generation resolves installation-known values before model execution, trading model-side binding work for setup, regeneration, and derived-copy maintenance
- Generation confidence does not by itself certify soundness (note) - Distinguishes next-token probability from factual truth and inferential validity: confidence can support correctness decisions only after task-specific validation, and high-assurance acceptance still needs a separate check
- Gödel machines are a proof-governed case of reflective self-modification (note) - A Gödel machine admits self-rewrites through proof under its current formalization; this restricts admission without establishing how many useful changes are reachable or how reliably they are found
- History has one chance to become checkable (note) - An artifact's production history is convertible to later-checkable form only at production time, via records/attestation or re-derivability; after that a bounded reviewer sees only carried state
- Holding a program theory means sustaining coherent search under delayed feedback (note) - Holding a program's theory is tested by whether a partial, fallible account of what the program is for keeps modification search, backtracking, and recovery coherent until delayed evidence arrives, not by whether the first change is right
- Human analogies can motivate functions without determining component boundaries (note) - Distinguishes functions and failure modes suggested by human cognition from the unsupported inference that an engineered agent should bundle the responsible roles along human boundaries.
- Human writing structures transfer to LLMs because failure modes overlap (note) - Writing genres evolved to prevent reasoning failures; the same structures help LLMs because they share those failure modes (content effects on reasoning) — evaluated per convention, not by analogy
- Human-LLM differences are load-bearing for knowledge system design (note) - Knowledge systems both inherit human-oriented materials and produce dual-audience documents (human + LLM), making human-LLM cognitive differences a first-class design concern rather than a generic disclaimer
- Improvements can accumulate without compounding (note) - Improvements accumulate when later improvement consumes or preserves retained results; compounding requires an earlier benefit to counterfactually improve a later episode, directly or through reinvested savings
- Improvements outside the admitted formal language need a pre-formal stage somewhere (note) - An improvement whose concepts have no expression in a loop's admitted formal language is reached only through a pre-formal stage, inside the loop or fixed at design time in the choice of language; translation relocates that stage
- In-context learning presupposes context engineering (note) - In-context learning only works when the right knowledge reaches the context window — the selection machinery that ensures this is itself learned and refined over deployment
- Inbound and outbound links serve asymmetric reader needs (note) - Outbound links are authored reader aids; their on-demand inverse serves distinct standing, grounding, impact, and tension needs without forbidding independently useful reciprocal links
- Increasing computational autonomy relocates human effort to the frontier instead of reducing it (note) - In an open-ended system, increasing computational autonomy need not cut total human hours — attention moves to the frontier — so measure improvements per human judgment, not human time
- Index completeness does not determine editorial orientation (structured-claim) - Complete generated listings establish membership, but their inputs do not determine topic-specific grouping, role phrases, or reading order without editorial judgment
- Indexes lower recall when they suppress retrieval that would find more (note) - A plausibly exhaustive index lowers route-level recall only when it prevents retrieval that would have produced greater task-relevant coverage
- Information value is observer-relative (note) - Information value is observer-relative: prior knowledge, tools, compute, and goals determine extractable structure, grounding use-shaped reshaping and discovery.
- Inspectable artifact, not supervision, defeats the blackbox problem (note) - Chollet frames agentic coding as ML producing blackbox codebases — codification counters this not by requiring human review but by choosing readable artifacts (code, prompts, schemas) that any agent can inspect, diff, test, and verify
- Instantiation alone cannot model agent learning across sessions (note) - The class/instance analogy captures session startup but omits the retained update relation that can revise later agent definitions and reusable-content placement
- Instruction specificity should match loading frequency (note) - The loading hierarchy (CLAUDE.md → skill descriptions → skill bodies → task docs) should match instruction specificity to loading frequency — always-loaded context competes for attention every session
- Instructions are typed callables with document type signatures (note) - Skills and tasks are typed callables — they accept document types as input and produce types as output, and should declare their signatures like functions declare parameter types.
- Intent controls a local choice only when it distinguishes its live alternatives (note) - A stated intent controls a local choice only when it changes which live alternatives are admissible, preferred, worth further search, or sufficient to stop
- Intent-framed delegation is a control regime; prompt length does not establish it (note) - For consequential agent handoffs, explains how Commonplace doctrine, task intent, binding bounds, and local evidence govern adaptive choice while prompt length does not.
- KB goals in always-loaded context guide inclusion decisions (note) - Without explicit goals in the always-loaded control-plane file, agents cannot reject well-written but off-scope material — a universal quality guide provides writing criteria but not domain scope
- KB maintenance (tag-readme) - Index of notes about keeping the KB healthy over time — detection of staleness and quality degradation, maintenance operations, and the dynamics that govern system entropy
- Knowledge storage does not imply contextual activation (note) - Separates knowledge that exists, knowledge loaded into context (read-back), and knowledge that actually changes behavior (activation); explains why retrieval and long context do not guarantee activation
- Knowledge-access architecture must be evaluated end to end, not by retrieval alone (note) - Explains why retrieval measures and storage-substrate labels cannot proxy for task-relative quality across discovery, loading, transformation, activation, and upkeep
- Known-target discovery benchmarks show reachability, not discovery closure (note) - Distinguishes backcast and reinvention benchmarks from autonomous discovery: they show that target insights are reachable from supplied ingredients, not that a system can select and verify new discoveries prospectively.
- Learning inside a fixed decomposition inherits its mistakes (note) - Why optimization cannot repair consequential distinctions, responses, or mappings outside the effective update space of a fixed task decomposition
- Learning is not only about generality (note) - Per Simon, any capacity change is learning; accumulation is the basic operation, explanatory-reach its key property (facts low, theories high); capacity splits into generality vs reliability/speed/cost
- Learning theory (tag-readme) - Curated head for the learning-theory tag — how systems learn, verify, and improve, with routes through its major child areas
- Legal drafting solves the same problem as context engineering (note) - Legal drafting parallels context engineering because both write ambiguous natural-language specifications for judgment-based interpreters, but law develops constraining more than codification
- Lightweight search control allocates further search without licensing adoption (note) - A search judgment is lightweight when its authority stops at allocating further investigation, probing, continuation, suspension, or abandonment rather than licensing an operative change
- Link graph plus timestamps enables make-like staleness detection (note) - Existing links already encode dependency information; comparing note and target timestamps flags notes that may be stale without any new annotation, analogous to make's file-based rebuild logic.
- Link strength is encoded in position and prose (note) - Not all links are equal — inline premise links ("since [X]") carry more weight than footer "related" links. Position and prose encode commitment level, creating a weighted graph that affects traversal, scoring, and quality signals.
- Link-following and search impose different metadata requirements (note) - Compares contextual local steps with long-range search and explains why these recurring navigation modes impose different metadata requirements on an agent knowledge base
- Linking theory (note) - Links are decision points; link quality is the reduction of navigation uncertainty per token of context consumed. Grounds our relationship vocabulary, title-as-claim, and position-encodes-strength practices under one model.
- Links (tag-readme) - Index of notes about linking — how links work as decision points, navigation modes, link contracts, and automated link management
- Links encode conditional possibilities, not obligations (note) - Links encode conditional possibilities, not obligations — every label must name a specific reader-need (the condition under which following pays off); content required for all reachable readers should be inlined, not linked
- Literature reuse can reverse a paper’s hierarchy of contributions (note) - Explains why conceptual distinctions built to support a paper's stated result can become its most valuable reusable output in a different research context
- LLM context is composed without scoping (note) - Flat context concatenation lacks local scope and produces name collision, contamination, and spooky action at a distance; code-built sub-agent contexts must impose boundaries
- LLM contexts interpret instructions and content through the same token medium (note) - LLMs interpret instructions and content through one token medium, enabling natural-language artifacts to alter behavior without translation while requiring architecture to enforce role, scope, and authority boundaries
- LLM debugging starts with retry-versus-rewrite triage (note) - Uses execution-versus-interpretation failure to choose the first debugging move: retry a bad execution of a sound reading, or rewrite a specification that reliably induces the wrong reading
- LLM frameworks should keep the tool loop optional (note) - Framework-owned tool loops package the common model/tool/retry pattern well, but strong frameworks keep the loop optional so applications can control state projection, branching, and re-entry
- LLM generation can hide a relaxed goal where human writing exposes a stall (note) - An LLM can ship fluent output after silently relaxing an unmet goal, while human composition may expose the same gap as a stall; a conjectural mechanism for why readers inherit the check
- LLM learning phases fall between human learning modes rather than mapping onto them (note) - Pre-training acquires both structural priors (evolution's role in humans) and world knowledge in one pass — making it and in-context learning intermediate on the evolution-to-reaction spectrum
- LLM output deviation requires three-way diagnosis because remedies target different relations (note) - For a fixed assembled input, whether V exceeds I, whether D escapes V, and how D's spread affects realization are three diagnostic questions with different primary repair surfaces
- LLM recompute cost shifts the store-vs-recompute balance (note) - For model-facing derived values, costly model-side recomputation shifts cache economics toward checked materialization, but persistence pays only when its total expected cost beats the alternatives and the copy substitutes for work
- LLM reliability (tag-readme) - Why LLM output deviates from intent — underspecification, interpreter failure, indeterminism — and the machinery for detecting and correcting it, from oracle theory and error correction to architectural separation
- LLM-executed methodologies are metacircular interpreters, not compilers (note) - Self-hosting LLM methodologies are closer to metacircular interpreters than compilers: agents re-interpret natural-language rules each session, while stable paths codify into validators and commands
- LLM-mediated schedulers are a degraded variant of the clean model (note) - When the agent scheduler lives inside an LLM conversation it becomes bounded and degrades; three recovery strategies — compaction, externalisation, factoring into code — restore the clean separation to increasing degrees
- LLM↔code boundaries are natural checkpoints (note) - At each LLM↔code transition both semantic underspecification and execution indeterminism collapse simultaneously, making these boundaries natural places to anchor debugging, testing, and refactoring
- Load-bearing vocabulary collisions should be prevented or visibly scoped at write time (note) - Unqualified technical senses have no reliable namespace in natural-language content; schema slots, rare compounds, and linked clause frames scope them at write time; audits and remediation recover when prevention fails
- Local materialization should outperform distant natural-language declarations (note) - Predicts that, for distant or non-obvious uses of a natural-language declaration, generated local materialization will outperform declaration-only presentation without creating a second maintenance authority
- Localized retention pays when sparse changes have bounded impact in a matching decomposition (note) - Addressable retention localizes a sparse change when units match its decomposition; total adaptation stays local only when the affected units also have a small, explicit impact closure
- Machinery persists by warrant, not position, in a reflective loop (note) - Reflection makes selected production machinery challengeable, but placement alone neither warrants nor requires revision; fixed general machinery may persist when its role and scope are earned
- Maintenance capacity must match harmful-artifact inflow (note) - Stable quality depends on capacity to prevent, contain, detect, and repair harmful retained artifacts keeping pace with their risk-weighted inflow, for which gross generation volume is only a proxy
- Maintenance operations catalogue should stage stable procedures for instructions (note) - Catalogue of periodic KB maintenance operations and readiness status, used as a staging ground before promotion into kb/instructions procedures
- MCP bundles stateless tools with a stateful runtime (note) - MCP forces stateless tool operations through a persistent server process — most tools are pure functions that don't need session state, connections, or lifecycle management, but pay the complexity tax anyway
- Measuring autonomy well enough to see it improve is an open problem (note) - Autonomy is reported per function rather than scored as a percentage, but that profile does not yet support comparison across systems or time
- Mechanistic constraints make Popperian KB recommendations actionable (note) - Bounded context and underspecification don't just permit conjecture-and-refutation — they require it; derives three concrete practices (falsifier blocks, contradiction-first connection, rejected-interpretation capture) from KB mechanics.
- Memory design adds operational axes to artifact analysis (note) - Memory design needs operational policy axes (capture, derivation, activation, authority assignment, lifecycle, evaluation) on top of substrate, form, lineage, and behavioral authority
- Memory-backed personalization can look like model improvement (note) - Distinguishes user-specific gains supplied by retained intent from gains in the model that interprets the assembled context.
- Methodological and computational closure track different changes (note) - Methodological closure tracks what a retained method settles; computational closure tracks the absence of human decisions during the assessed operation, without requiring every judgment to have explicit criteria
- Methodology enforcement is constraining (note) - Explains why enforcement strength is a partial order over activation and response semantics, rather than a fixed instruction-to-skill-to-hook-to-script ladder
- Methodology with incomplete coverage and its live theory fallback form a two-layer execution system (note) - In open or incompletely covered domains, the theory-derived fast path and live theory fallback co-execute while methodology-native content follows a separate maintenance regime
- Minimum viable vocabulary is the naming set that most reduces extraction cost for a bounded observer (note) - Defines minimum viable vocabulary as the names that most reduce a bounded observer's extraction cost, connecting conceptual thresholds to an information-theoretic optimization
- Mixed epistemic status must be preserved below the document level (note) - A document can combine observations, deductions, and plausible explanations; KB writing and review must retain which claims and transitions have which warrant.
- Model-resolved indirection adds interpretation work to LLM execution (note) - A reference adds model-side interpretation only when the model must resolve it; upstream literalization is worthwhile when binding, token, authority, and regeneration costs favor it
- Moving the interpretation–enforcement boundary requires cross-form coverage (note) - Moving responsibility between model-interpreted rules and formal enforcement crosses natural-language and symbolic forms, so governing the transfer requires coverage of both and their mapping
- Narrowing bought to survive review is paid for in content (note) - Repairing a defeated claim by shrinking its subject is justified at every step, but shrinking the subject into the predicate's own extension yields an analytic title that passes every gate and says nothing.
- Natural-language project state may specialize weight-resident search heuristics (note) - The natural-language part of project state may specialize general search heuristics already represented in an LLM's weights by supplying current intent, theory, branch history, and constraints
- Naur's compiler case tests one historically bounded documentation-and-consumption system (note) - Naur's compiler transfer failure rules out more documentation of the same kind, but tested one historically bounded package and consumption process rather than every possible rationale, indexing, retrieval, and activation system
- Naur's human-only conclusion needs more than the absence of explicit criteria (note) - Naur's human-only conclusion needs a further premise connecting unformulated judgment to computational inability; this reading preserves his functional tests without claiming that learned criteria are inexpressible
- Notes need quality scores to scale curation (note) - As the KB grows, /connect will retrieve too many candidates — evidence, type, inbound links, recency, and link strength can rank what is worth evaluating
- Observability (tag-readme) - Index of notes about making hidden state, hidden failure, and quality drift visible — runtime inspectability, degraded-execution signals, and maintenance-oriented detection mechanisms
- Opacity is a scale threshold, not a class property (note) - Opacity is not a representational form; any representation becomes practically opaque at sufficient scale, though distributed-parametric artifacts cross that threshold earliest.
- Open-domain memory retention needs a declared output spec (note) - Explains why an input stream alone can't answer 'what to store' in open-domain memory design; a declared output spec supplies the missing inclusion criterion.
- Open-ended construction builds an object and a theory of it (note) - Proposes that construction which must discover and revise an object's organization produces project-specific understanding beyond the object, using programs and theories as its two main cases
- Open-ended improvement must allocate search before decisive evaluation is available (note) - Open-ended improvement must choose which questions, candidates, experiments, or proof paths to develop before decisive evidence about them is available; even a Gödel machine's proof gate retains this prior search problem
- Open-ended theory learning and factory learning close the same reflective loop (note) - Derives one reflective loop from both open-ended theory learning and software-factory learning, and places the Gödel machine by transition licensing and theory provenance
- Operational signals that a component is a relaxing candidate (note) - Operational signals for when a component likely encodes a brittle proxy theory rather than an exact specification and should be relaxed instead of codified harder
- Opposed recompute factors do not decide documentation segmentation (note) - When audiences trade per-reconstruction savings against recurrence, neither factor alone determines cache value; among feasible alternatives, segmentation separately depends on whether specialization repays every cost introduced by the split
- Oracle accumulation improves selection for later candidates in its maintained domain (note) - A failure retained as a lesson helps tasks that retrieve it; retained as a maintained check it improves selection for later candidates in its domain and amortizes validation
- Oracle strength spectrum (note) - Exploratory framework — oracle strength, how cheaply correctness can be verified, as the gradient underlying the exact-spec/proxy-theory distinction, with an oracle-hardening pipeline
- Orchestration strategies and run-state have opposite persistence economics (note) - Separates ephemeral task-specific run state from reusable selection strategies inside host schedulers; RLM-style execution discards both and therefore loses the valuable reusable half
- Out-of-spec output is a failure of the interpreter, not the spec (note) - Interpreter failure is output that a spec's public meaning rules out; the fault attaches to the interpreter's role, so repair uses detection and correction rather than narrowing an already sufficient spec
- Parametric reproduction alone cannot replace an authoritative record (note) - Reproducing a record's content does not transfer its authority. Replacement requires a governed artifact with stable identity, integrity, contestability, and attribution; mutable records also require currentness and addressable revision.
- Periodic KB hygiene should be externally triggered, not embedded in routing (note) - Routing instructions serve the current task; periodic hygiene is triggered externally (user, heartbeat, CI), so embedding it in always-loaded routing blurs two responsibilities and adds session noise
- Pointer design tradeoffs in progressive disclosure (note) - Compares fixed, query-time, and crafted retrieval pointers across specificity, cost, availability, accuracy, and authoring dependence
- Preferential codification concentrates less predictable work at the agent boundary (note) - Explains the negative-selection mechanism by which preferential codification changes the composition of work retained at an agent boundary
- Problem matches guide method search; mechanism matches bound transfer (note) - Explains why problem matches generate source-method candidates rather than transfer warrant, why mechanism matches bound what transfers, and why composing bounded mechanisms requires target-side interaction checks.
- Process structure and output structure are independent levers (note) - Distinguishes constraints on reasoning steps from constraints on result shape and identifies the evidence needed to separate their effects
- Productive deferral requires a preserved option, discriminating evidence, and a convergence rule (note) - A qualitative test that separates controlled postponement of a consequential choice from hidden commitment, mere delay, and unmanaged indecision.
- Progressive constraining commits only after patterns stabilize (note) - Constraining via LLM code generation freezes a single projection of the spec in one shot, but progressive constraining observes behavior across many runs and commits only the interpretations that consistently emerge
- Project-theory possession requires comparing new demands with existing organization (note) - For open-ended modification, project-theory possession includes relating a new demand to existing responsibilities before parallel structure becomes the default; an explicit assimilation branch may counter additive coding-agent patches
- Promotion selects for unreliable activation, and the regress ends only at an external trigger (note) - Recasts promotion from 'the consumer lacks this' to 'the consumer will not apply this unprompted', and requires delivery to have a root firing event independent of that prior activation
- Prompt ablation converts human insight into deployable agent framing (note) - Methodology for testing prompt framings — uses controlled variation against a human-verified finding to identify which cognitive moves agents can reliably execute, then deploys the winning framing as instruction
- Psychology-to-agent transfer needs per-principle failure-mode testing (note) - Brainstorming a methodology for evaluating cognitive-science-to-agent transfer — assembled from three existing KB notes and tested against Youssef's five psychology principles as worked examples
- Quality signals for KB evaluation (note) - Catalogues graph-topology, content-proxy, and LLM-hybrid signals that could be combined into a weak composite oracle to drive a mutation-based KB learning loop without requiring usage data.
- Raw accumulation does not create usable memory (note) - Accumulation preserves material, but usable agent memory requires ingress work that adds handles, scope, relationships, provenance, trust signals, and lifecycle pressure.
- Reasoning production is not reasoning evaluation (note) - Review and critique systems need independent process-validity checks because a model can substitute answer reconstruction for reasoning evaluation
- Recognition, not linking, is the hard problem in knowledge systems (note) - Seeing that two items share a structure is the expensive step in connecting knowledge; articulating a seen connection is cheap, and naming a recognized structure amortizes later recognition.
- Reflection buys addressability (note) - Self-improvement can accumulate without reflection — parametric learners do — but non-reflective retention gives only indirect handles; reflective retention makes the changed object addressable
- Reflection makes retained lessons second-order: a lesson can reject or rescope a prior commitment (note) - Reflection lets a retained lesson target a prior commitment explicitly — rejecting, revising, or rescoping it — while non-reflective correction acts indirectly through the substrate
- Reflective coverage is graded across representational forms (note) - Reflective coverage is stated per represented form and operation profile; control of an external dependency does not make that dependency part of the system's reflective coverage
- Reflective theory refinement has separate structural, epistemic, and implementation lineages (note) - No single predecessor is closest to reflective theory refinement: runtime self-modeling supplies the self-target, classical theory refinement the mechanism with different fillers, and Workspace Optimization only an implementation analogy
- Reflective theory refinement needs interpretation, retention, and independent read-back (note) - Reflective theory refinement requires semantic interpretation, addressable retention, independent outcome read-back, and continuation on one causally co-indexed path; these are distinct functions that need not share one substrate
- Reliability dimensions map to oracle-hardening stages (note) - The four reliability dimensions from Rabanser et al. (consistency, robustness, predictability, safety) each harden a different oracle question — mapping empirical agent evaluation onto the oracle-strength spectrum
- Retained system-definition artifacts enable persistent deployment-time adaptation (note) - Retaining evaluated changes to behavior-shaping prompts, rules, tools, and tests gives deployed systems a persistent adaptation path outside model-weight updates
- Retaining episode evidence keeps a distilled rule open to re-examination (note) - Keeping relevant episode evidence and its relation to a distilled rule preserves a route for re-examining that rule; reconstruction, comparative value, and correct generalization still require testing
- Reverse compression is when LLM output expands without adding information (note) - LLMs can inflate compact seeds into verbose artifacts without adding extractable structure; a KB resists this only when links make additional structure accessible
- Review automation should target verifiable subroles before reviewer identity (note) - Scholarly-review automation should decompose reviewer work into separately verifiable subroles before giving an AI system reviewer-level authority
- Revising an improvement objective is licensed from outside it or is not improvement (note) - Objective change is improvement only against a level outside both objectives; proxy revision, re-indexing, and surfaced under-specification subtract most apparent cases
- Revision guided by rationale needs faithfulness, not just legibility (note) - When revision relies on a rationale to locate a failed premise, a misleading rationale can direct repair to the wrong part; retained rationale is optional for theory refinement
- RLM has the model write ephemeral orchestrators over sub-agents (note) - RLM packs orchestration over sub-agents into the tool-loop model by having the model write orchestrators in a REPL — elegant but ephemeral because the orchestrators are discarded after each run
- RLM, λ-RLM, Tendril, and llm-do separate restriction from persistence (note) - RLM variants, Tendril, and llm-do show that control-language restriction and artifact persistence are separate questions, including where cited RLM sources leave post-return lifecycle unspecified
- Rule-based context selection needs a pre-existing signal (note) - A rule-based selector can target one case only when a rule-ready signal already distinguishes it; otherwise the system must wait, load broadly, or infer relevance from task and candidate content
- Runtime structure determines the control surfaces available to governance (note) - Runtime structure and runtime governance are separable, but the runtime's structure determines which inspection, validation, correction, and drift-control operations governance can actually perform
- Scaling absorbs scaffolding at fixed task difficulty, not at the deployment frontier (note) - Stronger models shrink the scaffolding a fixed task needs; durable deployment-specific structure recurs at the frontier only while assigned difficulty keeps pace with capability and some reliability function stays advantageous to externalize
- Scenario decomposition drives architecture (note) - Deriving architectural requirements by decomposing concrete user stories into step-by-step context needs — not from abstract read/write operations but from what the agent actually has to load at each stage
- Scheduler-LLM separation exploits an error-correction asymmetry (note) - Symbolic bookkeeping eliminates underspecification, indeterminism, and bias relative to the implemented transition function; semantic work faces all three. Mixing forces exact state onto an expensive substrate; codification renegotiates the boundary
- Selecting an LLM output fixes a result, not its interpretation (note) - Selecting one LLM output for operative reuse creates a stable artifact-testing target without resolving ambiguity inside the text, so generator and artifact tests answer different questions
- Self-improvement is relative to a declared objective (note) - The improvement objective is a declared parameter alongside boundary and horizon, carrying two separable conditions — indexed by the analyst, antecedent in the pathway — whose failures differ in kind
- Self-improving systems (tag-readme) - Curated head for the self-improving-systems tag — membership, update architecture, and the four-part pathway profile; selective picks
- Semantic review catches content errors that structural validation cannot (note) - Structural validation catches form errors; semantic review catches content errors like incomplete enumerations, grounding drift, boundary-case gaps, and internal contradictions
- Semantic sub-goals that exceed one context window become scheduling problems (note) - Some semantic subgoals exceed one context window, so they must be partitioned into smaller semantic judgments with symbolic collection, filtering, and staged summarization between them
- Semantic work can be relocated but not eliminated (note) - A meaning-dependent judgment is never removed, only placed — moved upstream where its inputs exist (amortized) or off a bottlenecked context (offloaded); 'free at use-time' always means paid earlier
- Session history should not be the default next context (note) - Storing execution history and loading it into the next agent call are separate decisions; chat and framework-owned tool loops conflate them by making session history the default next context
- Short composable notes maximize combinatorial discovery (note) - The library's purpose is to produce notes that can be co-loaded for combinatorial discovery — short atomic notes are a consequence of this goal; longer synthesized artifacts belong in workshops or derived instructions
- Silent disambiguation is the semantic analogue of tool fallback (note) - When an agent silently resolves unacknowledged material ambiguity in a spec, final success hides that the contract failed to determine the path — an extension of the tool-fallback observability problem
- Skill discovery re-fires in every sub-agent context, not just the top-level invocation (structured-claim) - Skill discovery is per-context and autonomous — every installed skill is re-matched in each sub-agent context, even ones a parent narrowed, so a delegating skill's own discoverability is a leak vector
- Skills are instructions plus routing and execution policy (note) - Skills add structured discovery, user-facing invocation, and declarative execution policy (tool permissions, model override, context isolation) beyond the shared procedure
- Skills derive from methodology (structured-claim) - The methodology→skill relationship is same-medium derivation, with the methodology retained as live fallback — distinct from codification and constraining
- Soft degradation can bind before the hard cap even when required evidence fits (note) - For quality-sensitive agent work whose required evidence fits within the provider window, volume, complexity, and interference can silently constrain usable context before the hard cap
- Soft-bound traditions as sources for context engineering strategies (note) - Survey of twelve soft-bound traditions as candidate sources for context engineering strategies, with a three-tier assessment of what transfers, what's plausible, and what's blocked
- Solve low-degree-of-freedom subproblems first to avoid blocking better designs (note) - Ordering heuristic for decomposition: commit first to decisions with the fewest viable options, then place flexible choices around them to preserve global optionality.
- Source changes should surface downstream review targets, while reverse lineage can remain searchable (note) - Source-dependent artifacts need lineage signals when an upstream change may render those artifacts stale, regardless of where the lineage record is stored
- Spec mining is codification's operational mechanism (note) - Operationalizes codification by extracting deterministic verifiers from observed stochastic behavior — the mechanism that converts blurry-zone components into calculators
- Specification strategy should follow where understanding lives (note) - Among durable artifacts, spec-first, bidirectional spec, and spec mining fit different phases: when understanding is available upfront, discovered during execution, or only visible after observation
- Specification-level separation recovers scoping before it recovers error correction (note) - OpenProse-like DSLs expose control flow and discretion boundaries while leaving scheduling and validation on the LLM substrate, creating an intermediate regime between flat prompting and symbolic scheduling
- Stale self-description conceals its own staleness (note) - What artifact drift adds when it is reflexive: the process that would detect it consults the artifact that drifted, the trigger has no edit event to hook, and synchronization load scales with autonomy
- Stateful tools recover control by becoming hidden schedulers (note) - Granting the strongest stateful-tool escape hatch shows that recovered control comes from relocating the scheduler into an exceptional tool or runtime, not from the framework loop itself
- Structured output is easier for humans to review (note) - Separated Evidence and Reasoning sections let human reviewers check facts and logic independently — a purely readability argument that doesn't depend on LLM behavior at all
- Structured-prompt gains do not establish training-distribution selection (note) - Formatting compliance, extra computation, and task decomposition can mimic distribution-selection gains, so prompt performance alone cannot identify the mechanism
- Subtasks that need different tools force loop exposure in agent frameworks (note) - When decomposition creates child tasks with different tool surfaces, the parent must construct fresh calls for each child, so a framework-owned loop is no longer the right control surface
- Superseded choices need a historical witness; refuted beliefs lose subject-matter standing (note) - Retention after supersession follows remaining truth role rather than maintenance operation: preserve a witness to the choice event, while a refuted belief loses subject-matter standing
- Synthesis is not error correction (note) - Synthesis propagates errors by merging all agent outputs; voting corrects errors by discarding minorities — Kim et al.'s 17.2× amplification is a synthesis failure, not evidence against multi-agent coordination
- System use is an initial selection environment when theory fit lacks a fixed oracle (note) - When no complete fixed oracle decides whether a claim belongs in a working theory, distributed consequences of live system use can provide an initial selection environment
- System use provides evidence of theory fit and causal usefulness, not independent warrant (note) - Consequences of using a claim in a live system can test its integration and causal usefulness, but independent factual, formal, source, or scope evidence is still needed for its warrant
- System-definition artifacts are crystallized reasoning under context scarcity (note) - Separates heuristic rules that substitute for unavailable read-time reasoning from authority-bearing constraints and symbolic codification, which remain useful even with abundant context
- Systematic prompt variation serves verification and diagnosis, not explanatory-reach testing (note) - Controlled prompt variation either decorrelates checks or measures brittleness under fixed task semantics; Deutsch's variation test instead changes the explanation to test mechanism and explanatory-reach
- Tags (tag-readme) - Hub for all tag READMEs — browse the KB by conceptual domain rather than by directory; complete over the tag pages in this collection
- Task families and product families classify different things (note) - Task families group obligations or evaluations; software product families group products through declared commonality, variability, and reusable production scope
- Technical constraints turn KB objective-function choice from philosophy into engineering (note) - Three technical constraints and the codification lever make KB objective-function choice testable engineering, not philosophy; goals set the loss, local contracts specialize it, and oracle strength differs by objective
- Text testing framework — source material
- The augmentation-automation boundary is discrimination not accuracy (note) - Crossing from augmentation to automation requires per-instance discrimination, not aggregate accuracy — discrimination is empirically stagnant, so scaling capability alone cannot cross the boundary
- The Bitter Lesson defense portfolio has one load-bearing member for the form-only rebuttal (note) - The production-method versus representational-form distinction answers only a narrow weights-only inference; theory-guided bootstrapping is a provisional first strategy under incomplete global evaluation, not a defense of continuing hand production
- The bitter lesson selects against unearned reach, not against structure (note) - The lesson selects against claims whose reach was asserted rather than earned by a refuting test, not against structure or origin — theory search in readable forms is its own method; earned reach protects the claim, not its carrier
- The bitter lesson selects production methods, not representational forms (note) - The lesson's axis is production method — hand-crafted versus search-and-learning — not representational form. Learned localized forms are therefore a coherent scaling hypothesis, with cross-artifact credit assignment as the decisive open problem
- The boundary of automation is the boundary of verification (note) - Synthesis — oracle theory, labor economics, frontier-lab capability predictions, and supply-chain integrity evidence converge on verification cost as the primary structural determinant of automation
- The chat-history model trades context efficiency for implementation simplicity (note) - Chat history persists because appending messages preserves information and avoids interface design, but that convenience trades away selective loading under bounded context
- The deployed system, not the model alone, is the unit of learning (note) - Because prompts, retrieval, tools, and runtime policy jointly determine deployed behavior, model-only learning leaves consequential system choices fixed
- The four-field record exposes an efficiency, security, and sovereignty risk triad (note) - The four artifact-analysis fields exist to surface three architectural review concerns over retained behavior — efficiency, security, and sovereignty — with sovereignty (owner control to inspect, regenerate, delete, roll back) as the new axis
- The practical scheduler is the host language, not a reified select (note) - The simplest practical orchestration library demotes the tool loop to a returning, per-call-parameterized function and lets ordinary host-language code play select and K — reifying K only when the run must outlive its process or outgrow its memory
- The readable-artifact loop is the tractable unit for continual learning (note) - Identifies the natural-language-plus-symbolic pair as the tractable first loop for representational-form coevolution because it shares context, operates at current tempos, and already has a codification boundary
- The self-improving-system definition classifies its boundary cases without ad hoc exceptions (note) - Ten boundary cases run against the self-improving-system definition — each classifies by the stated criteria alone; the stress they apply falls on boundary declaration, not on the membership clauses
- The verifiability gradient (note) - Symbolic artifacts sit on a gradient from loose natural-language to deterministic code; higher-verifiability artifacts support tighter iteration loops, and learning moves artifacts along it in both directions
- The wikiwiki principle: lowest-friction capture, then progressive refinement in place (note) - Ward Cunningham's wiki design principle — minimize capture friction, refine in place — drives the text→note→structured-claim codification ladder
- Theory building and capacity building make the same kind of fallible commitment (note) - Theory building and capacity building both retain resolutions their evidence does not entail; an explanatory commitment stays answerable to the object it describes while a constructive commitment changes the object, so retraction differs in kind
- Theory mediation can coordinate heterogeneous factory development (note) - Fallible natural-language project theory may provide an addressable way to coordinate heterogeneous factory development while search, testing, and backtracking construct and revise it
- Theory refinement may improve sample efficiency under structured shifts (note) - Conjecture: theory refinement, learning that discovers, assesses, and revises addressable theories, may need fewer target observations when a shift preserves the structure a theory names
- Theory warrant should be tracked at the finest granularity evidence licenses (note) - Treat support for a theory as warrant for only the most specific claim, conjunction, model, and scope the evidence identifies; do not distribute joint warrant beyond what it entails without additional attribution
- Three-space agent memory echoes Tulving's taxonomy but the analogy may be decorative (note) - The value of separating knowledge, self, and operational memory is that each has a different lifecycle — accumulation, slow evolution, and high churn; whether the Tulving mapping adds explanatory power beyond different retention policies is open
- Title as claim enables traversal as reasoning (note) - When note titles are claims rather than topics, following links between them reads as a chain of reasoning — the file tree becomes a scan of arguments, and link semantics (since, because, but) encode relationship types
- Title as claim exposes commitments, enabling Popperian maintenance (note) - When an index is a list of claims rather than topics, reviewing the KB becomes scanning hypotheses — each title exposes its commitment and invites the question "do I still believe this?" without opening the file
- Title as claim makes overlap between notes visible (note) - When note titles are claims, overlap between notes is visible at the index level — similar assertions are obvious without opening files; topical titles hide overlap behind different labels for the same territory
- Tool loop (tag-readme) - Index for the tool-loop argument — the framework-owned tool loop is useful but should yield control when tasks need different tool surfaces, exceed one context window, or codify scheduling
- Tool usefulness, computational autonomy, warrant, and system power are separate dimensions (note) - Tool usefulness, computational autonomy, warrant, and system power move independently in a human-agent system, so a progress claim has to say which one moved and autonomy gains do not license power claims
- Topology, isolation, and verification form a causal chain for reliable agent scaling (note) - Topology, isolation, and verification may form a strict dependency chain rather than independent design choices — tested against the simpler account that good decomposition implies the other two
- Trace-extracted memory earns authority per operation, not at capture (note) - Trace memories begin as records; verification, abstraction, and consultation earn authority under progressively harder oracles, while unverified stores accumulate guesses presented as knowledge
- Traditional software can bracket executor conformance; LLM systems cannot (note) - Wrongness is a relation to a norm, never intrinsic to a computation; classical stacks bracket the executor-conformance norm so every failure resolves to the spec, and LLM systems cannot, which is what generates the three-source deviation taxonomy
- Traversal improvements should be deferred via logging to avoid mid-task context switching (note) - Loading writing methodology into an already-committed context window is expensive; a one-line log entry preserves the improvement signal at near-zero cost and lets a separate pass do the fix
- Treat continual learning as representational-form coevolution (note) - Behaviour change spans distributed-parametric, natural-language, and symbolic forms, so the question is how their improvement loops relate — not which is the real locus of learning
- Two context boundaries govern collection operations (note) - Distinguishes the body-loading boundary from the later title-and-description index boundary, yielding three collection-size regimes with different consequences for areas, connect, and whole-KB work
- Type system (tag-readme) - Index of notes about the document type system — why types exist, what roles they serve, how they improve output quality, and how they're structured
- Type system enforces metadata that navigation depends on (note) - Descriptions don't appear spontaneously — they exist because the note base type requires them; without enforcement, metadata degrades and navigation collapses to opening every document
- Types give agents structural hints before opening documents (note) - Types and descriptions let agents make routing decisions without loading full documents — the type says what operations a document affords, the description filters among instances of that type
- Under sub-agent decomposition, feasibility is the heaviest fork's net load (note) - Shows why decomposition changes feasibility from total operation cost to the largest residual load left on any fork after work is shifted to siblings or the parent
- Underspecification and indeterminism complicate programming for prompts in distinct ways (note) - Indeterminism doubles test runs (statistical testing over distributions); underspecification doubles test targets (spec analysis for ambiguity). Conflating the two leads to misdiagnosis
- Unified calling conventions enable bidirectional refactoring between neural and symbolic (note) - When agents and tools share a calling convention, components can move between neural and symbolic without changing call sites — llm-do demonstrates this with name-based dispatch over a hybrid VM
- Unit testing LLM instructions requires mocking the tool boundary (note) - Skills are programs whose I/O boundary is tool calls — mocking that boundary creates controlled environments for testing whether instructions produce correct behavior, complementing text artifact testing with instruction-level regression detection
- Universal software factory needs a declared universality axis (note) - Universal software factory is ambiguous unless the universality axis, covered class, supplied inputs, adequacy relation, and resource bounds are declared
- Use tests a decomposition locally; retained rationale is what makes transfer testable (note) - Running a decomposition confirms only that it sufficed here; because many force-sets fit the same split, rationale retained at design time is what gives a transfer claim an antecedent to test
- Verification needs a typed target before it needs an oracle (note) - A check's warrant depends on a declared target class, so an unverifiable heterogeneous layer is usually blocked by missing artifact classification, not oracle difficulty — ontology precedes oracle
- Vibe-noting (note) - Linked, maintained knowledge artifacts let LLM agents recover reasoning across sessions, improving augmentation even when weak verification still blocks automation
- Warranted autonomy is bounded by oracle domain (note) - Bare autonomy is free, but warranted evaluation autonomy extends only to the candidates an oracle can assess with the required confidence
- Warranted reader update is the objective of substantive writing (note) - Defines epistemic interestingness as a relevant, warranted change relative to an intended reader's prior, making contribution selection—not accumulated inputs—the purpose of multistage writing.
- Warranted transfer out of the human cut leaves people the hardest-to-warrant decisions (note) - When a system preferentially transfers decisions whose premises, criteria, and checks are available, the remaining human decisions become harder to warrant per decision; this predicts a residue composition, not structural computational openness
- Weakly discriminated qualities tend to be underselected (note) - Statistical conjecture: under named proposal-selection conditions, unequal oracle discrimination yields unequal enrichment; absolute degradation needs an additional directional mechanism
- Weight-resident methodologies provide context-efficient behavioral compression (note) - A compact cue can activate a much larger methodology already represented in model weights, trading very low context cost for model-dependent reconstruction rather than exact retained specification
- Why directories despite their costs (note) - Directories buy one–two orders of magnitude of human-navigable scale over flat files, and enable local conventions per subsystem — but each new directory taxes routing, search config, skills, and cross-directory linking
- Why notes have types (note) - Seven roles of the type system — navigation hints, metadata enforcement, verifiable structure, local extensibility, content-layer identification, output quality through structured writing discipline, and maturation through constraining
- World models assess explanatory-reach through action-conditioned prediction (note) - Learned world models can assess explanatory-reach when action-conditioned predictions are tested across the interventions or shifts a commitment claims
- Writing conventions for kb/notes/
- Writing styles are strategies for managing underspecification (note) - Maps descriptive, prescriptive, prohibitive, explanatory, and conditional context-file styles to distinct ways of narrowing agent interpretation, each trading constraint against generality