continual-learning
Type: types/tag-readme.md
Assign this tag to work on continued learning through retained changes outside model weights: prompts, rules, tools, schemas, and tests; how those changes are governed; and how this learning relates to updates in other representational forms. Weight-only training and post-release change without an account of learning through retained artifacts are insufficient. Two notes establish the tag: Retained system-definition artifacts enable persistent deployment-time adaptation and The readable-artifact loop is the tractable unit for continual learning. Bitter Lesson notes defend the bet that learning in readable forms can scale. It is a child of self-improving-systems. Boundary: deploy-time-learning is the phenomenon, that deployment reveals what design could not; this tag holds the mechanisms that answer it.
Where learning lives in a deployed system
- Retained system-definition artifacts enable persistent deployment-time adaptation — the framework claim: evaluated artifact changes persist outside weights
- The deployed system, not the model alone, is the unit of learning — prompts, retrieval, tools, and runtime policy jointly set behavior, so model-only learning leaves them fixed
- Constraining during deployment is continuous learning — narrowing interpretations through prompts, schemas, and tests is one form the accumulation takes
- Instantiation alone cannot model agent learning across sessions — the class/instance picture misses the update relation across sessions
- Discarding all experience-dependent state prevents cross-run accumulation — the failure condition: nothing learned survives when no experience-dependent state carries over
- Factory learning is experience-responsive retention that improves the factory — the same mechanism applied to reusable production machinery
Loops across representational forms
- Treat continual learning as representational-form coevolution — parametric, natural-language, and symbolic forms each change; the question is how their loops relate
- The readable-artifact loop is the tractable unit for continual learning — start with the natural-language-plus-symbolic pair, which shares context and a codification boundary
- LLM-executed methodologies are metacircular interpreters, not compilers — rules are re-interpreted every session while stable paths codify into validators and commands
- Moving the interpretation–enforcement boundary requires cross-form coverage — shifting a rule into code crosses forms, so governing it needs coverage of both and their mapping
- Improvements outside the admitted formal language need a pre-formal stage somewhere — a formal-only loop relocates pre-formal work, it does not remove it
Governing updates and their limits
- Continual learning requires governing behaviour-changing writes, not just storing content — persistence is not enough: updates must be selected, validated, authorized, and coordinated
- Learning inside a fixed decomposition inherits its mistakes — optimization cannot repair distinctions outside the decomposition's update space
- An optimal long-run learning strategy invests in its own machinery — a machinery improvement is reused by every later episode, so it can out-earn immediate learning
- Automating KB learning is an open problem — the KB learns through manual improvement; automating judgment-heavy mutations lacks oracles
The Bitter Lesson defense
- The bitter lesson selects production methods, not representational forms — the lesson's axis is hand-crafted versus search-and-learning, so learned readable forms remain a coherent scaling bet
- The bitter lesson selects against unearned reach, not against structure — what the lesson penalizes is asserted reach, not structure or origin
- A hand-crafted bootstrap fits the Bitter Lesson only if learning can outgrow it — the condition on a hand-built starting state: scalable learning must displace what it supplies
- The Bitter Lesson defense portfolio has one load-bearing member for the form-only rebuttal — which defense actually carries weight, and why bootstrapping is only provisional
Related Tags
- self-improving-systems — the parent; continual learning is the deployed-system case of self-improvement
- deploy-time-learning — the phenomenon these mechanisms answer; several members carry both tags
- learning-theory — the general account of learning these notes apply
- reflection — whether the retained changes are represented and selectively revisable
- warranted-autonomy — who may authorize the behavior-changing writes
- improvement-loop — the search, evaluation, and retention loop each form's learning runs through
- software-factory — shares the factory-learning note; continual learning applied to production machinery
- theory-builder — learning through retained, criticized theory