continual-learning
Type: types/tag-readme.md
Assign this tag to work on continued learning through retained changes outside model weights: prompts, rules, tools, schemas, and tests; how those changes are governed; and how this learning relates to updates in other representational forms. Weight-only training and post-release change without an account of learning through retained artifacts are insufficient. Two notes establish the tag: Retained system-definition artifacts enable persistent deployment-time adaptation and The readable-artifact loop is the tractable unit for continual learning. Bitter Lesson notes defend the bet that learning in readable forms can scale. It is a child of self-improving-systems. Boundary: deploy-time-learning covers what deployed use reveals and responses to it, including human maintenance. This tag requires learning through retained changes outside weights; work on that learning as a response to deployed experience may carry both tags.
The distinction from improvement-loop is how learning persists outside model weights versus how changes are selected. Assign this tag when evidence becomes a retained change used in later operation, or the note develops that path's governance, limits, or relation to weight updates. Saving a file or issuing a one-off prompt is insufficient. A separate candidate-admission gate is not required; a note that also develops proposal selection can carry both tags.
Where learning lives in a deployed system
- Retain reusable orchestration strategies — selected control logic persists as tested library code, while task-specific state stays temporary
- A retained instruction preserves what testing selected — an instruction retains an empirically selected procedure outside model weights
- Retained system-definition artifacts enable persistent deployment-time adaptation — the framework claim: evaluated artifact changes persist outside weights
- The deployed system, not the model alone, is the unit of learning — prompts, retrieval, tools, and runtime policy jointly set behavior, so model-only learning leaves them fixed
- Constraining during deployment is continuous learning — narrowing interpretations through prompts, schemas, and tests is one form the accumulation takes
- Instantiation alone cannot model agent learning across sessions — the class/instance picture misses the update relation across sessions
- Discarding all experience-dependent state prevents cross-run accumulation — the failure condition: nothing learned survives when no experience-dependent state carries over
- Factory learning is experience-responsive retention that improves the factory — the same mechanism applied to reusable production machinery
Loops across representational forms
- Treat continual learning as representational-form coevolution — parametric, natural-language, and symbolic forms each change; the question is how their loops relate
- The readable-artifact loop is the tractable unit for continual learning — start with the natural-language-plus-symbolic pair, which shares context and a codification boundary
- LLM-executed methodologies are metacircular interpreters, not compilers — rules are re-interpreted every session while stable paths codify into validators and commands
- Moving the interpretation–enforcement boundary requires cross-form coverage — shifting a rule into code crosses forms, so governing it needs coverage of both and their mapping
- Improvements outside the admitted formal language need a pre-formal stage somewhere — a formal-only loop relocates pre-formal work, it does not remove it
Governing updates and their limits
- Continual learning requires governing behaviour-changing writes, not just storing content — persistence is not enough: updates must be selected, validated, authorized, and coordinated
- Learning inside a fixed decomposition inherits its mistakes — optimization cannot repair distinctions outside the decomposition's update space
- Automating KB learning is an open problem — the KB learns through manual improvement; automating judgment-heavy mutations lacks oracles
The Bitter Lesson defense
- The bitter lesson selects production methods, not representational forms — the lesson's axis is hand-crafted versus search-and-learning, so learned readable forms remain a coherent scaling bet
- The bitter lesson selects against unearned reach, not against structure — what the lesson penalizes is asserted reach, not structure or origin
- A hand-crafted bootstrap fits the Bitter Lesson only if learning can outgrow it — the condition on a hand-built starting state: scalable learning must displace what it supplies
- The Bitter Lesson defense portfolio has one load-bearing member for the form-only rebuttal — which defense actually carries weight, and why bootstrapping is only provisional
Related Tags
- self-improving-systems — the parent; continual learning is the deployed-system case of self-improvement
- deploy-time-learning — what deployed use reveals and how systems respond; retained learning is one response, so several members carry both tags
- learning-theory — the general account of learning these notes apply
- reflection — whether the retained changes are represented and selectively revisable
- warranted-autonomy — who may authorize the behavior-changing writes
- improvement-loop — candidate search, reject-capable evaluation, and retention; one update architecture, distinct from direct updates
- software-factory — shares the factory-learning note; continual learning applied to production machinery
- theory-builder — learning through retained, criticized theory
Other tagged notes
- Ad hoc prompts extend the system without schema changes - Any system with an LLM agent layer can absorb new requirements through natural language prompts without changing the deterministic base
- Codification and relaxing navigate the bitter lesson boundary - Since you can't identify which side of the bitter lesson boundary you're on until scale tests it, practical systems must codify and relax — with spec mining avoiding the vision-feature failure mode
- In-context learning presupposes context engineering - In-context learning only works when the right knowledge reaches the context window — the selection machinery that ensures this is itself learned and refined over deployment
- Localized retention pays when sparse changes have bounded impact in a matching decomposition - Addressable retention localizes a sparse change when units match its decomposition; total adaptation stays local only when the affected units also have a small, explicit impact closure
- Promote Only When Future Value Exceeds Maintenance Cost - Candidate memory should become durable only when future retrieval or activation value exceeds review and maintenance cost
- RLM, λ-RLM, Tendril, and llm-do separate restriction from persistence - RLM variants, Tendril, and llm-do show that control-language restriction and artifact persistence are separate questions, including where cited RLM sources leave post-return lifecycle unspecified
- Spec mining is codification's operational mechanism - Operationalizes codification by extracting deterministic verifiers from observed stochastic behavior — the mechanism that converts blurry-zone components into calculators
- Use Trace Extraction As Meta-Learning - Trace extraction is an after-the-fact learning path that must respect signal quality, review, and readable-artifact versus distributed-parametric learning boundaries