Articles
Writing from the Commonplace knowledge base for readers outside the project — researchers and builders of agent and knowledge systems. Each article stands on its own, carries a byline and lifecycle status, and links into the KB it was distilled from.
Articles circulate in three states. A draft may change without notice and carries a visible banner inviting comments. A working paper is revisable on the record: its claims are still open, it carries a version and a revision date, and it invites counterexamples. A published article is frozen — corrections appear as dated annotations or superseding articles, never as silent rewrites.
Working papers
Nothing circulating yet.
Published
Nothing published yet.
Presentations
- Where It Lives Is Not What It Is — ASISAS 2026 slides — HTML deck prepared for a 10-minute talk and 2 minutes of questions, based on the June 23 position paper. Use arrow keys, space, or click to navigate.
The paper is published as Zbigniew Łukasiak, "Where It Lives Is Not What It Is: An Architectural Vocabulary for Retained Adaptation in Agentic Systems", in Software Architecture. ECSA 2026 Tracks and Workshops, Lecture Notes in Computer Science, Springer, 2026, pp. 286–295, doi:10.1007/978-3-032-39143-8_24.
The deck uses memory for the audience-facing term and natural-language in place of the paper's prose. It derives the three representational forms from assigned consequences and localization; the terminology and derivation are explained in representational form. The slides show the terminology changes with strikeouts.
In draft
Five drafts make one series: a lead article and four supplements. The lead states the idea and stands on its own. Each supplement develops one part of it, so read the lead first and then whichever supplement answers your next question. The series proposes Commonplace's experiments; it reports no completed experimental result from that program. The survey discusses results reported by other systems.
- Can a Theory Builder Running on Fixed-Weight LLMs Learn? — the lead. It introduces the theory builder, a system that runs Popper's cycle of problem, tentative solution, error elimination, and revised problem over explicit theories, defined by four conditions: stated theories, acted on, criticized for what they say, with the result shaping the next round. It states the bet that a fully automated theory builder that learns can be built with today's fixed-weight LLMs, argues that an LLM can apply precise definitions where formal proof cannot reach, gives the payoffs and costs of explicit learned state and three conjectured advantages, argues compatibility with the Bitter Lesson, states the two-part test that would count as learning, and says how Commonplace builds one: reflective from the start, moving operations from people to computation, on the conditional case that investing in the method pays over a long horizon. Testing, bootstrapping from Commonplace, the software-house arrangement, and the survey each point to a supplement.
The supplements, in the order a reader is likely to want them:
- Testing Whether a Theory Builder Learns — sets a modest first testing goal: show under controlled conditions that retained revisions causally improve later capacity and stay revisable. Gives seven tests that expose the causal role of retained state (retention, use, withholding, perturbation, reconstruction, transfer, revision), separates task success from learning, sketches controlled task families in which the revised builder is run against its declared seed on fresh instances whose answers are fixed outside it, with a declared evidence interface and controls, records what people contribute operation by operation, defers long-horizon accumulation and compounding, and quotes the three adopted whole-program hypotheses with their refuters. No run has been performed.
- Bootstrapping an Autonomous Theory Builder with Commonplace — develops the lead's approach to building its theory builder. Commonplace starts reflective, with an operator performing many of its operations; it retains what is learned, lets the builder build the software its knowledge needs, and moves operations to computation one at a time, recording what stays with the operator. Success is measured by the human decisions each completed, verified improvement requires.
- An Automated Software House as a Second Test of a Theory Builder — a companion arrangement to the lead's knowledge base. It states the conjecture that an automated software house capable of open-ended coherent change can operate practically with models available by 2026-09-02 and held fixed; argues that a house which states, criticizes, and revises its program theory (Naur's term) is a theory builder, and how that squares with Naur's view that the theory cannot be fully written down; explains why software supplies a stronger falsifier than a knowledge base at the price of a harder claim; and defines four conditions a witness run must meet, each with what would fail it.
- Which Existing Self-Improving Systems Are Theory Builders — compares eighteen systems, starting with reported tests of retained knowledge and skill improvement. Places each against the four theory-builder conditions and judges reported learning separately: five builders (two with reported gains, three human-staffed without a measured gain), one outside, and twelve unsettled, mostly because the evidence shows an outcome gate rather than criticism of what a retained unit says; no placement turns on how far results persist. Then proposes comparisons that would settle the placements and isolate what criticism contributes.
One draft stands outside the series:
- What an Automated Reviewer Should Measure — a proposal for automated review of explanatory articles, especially theory revisions. The reviewer assesses explanatory reach and explanatory economy, relative to a declared reader's prior account, as evidence of compression progress: capturing more of a domain's structure, or the same structure more simply. It gives the reviewer's procedure: recover five parts of a revision, derive consequences beyond the motivating cases, check them, and record the extra assumptions each extension needed. That leaves two claims to test separately: whether the assessment tracks compression progress, measured by predictive coding where that is feasible, and whether that progress improves the reader's performance on relevant tasks. A favourable assessment does not certify an article's evidence or reasoning. Nothing in it has been run.
Superseded
Nine earlier drafts have been withdrawn. Each address redirects to the draft that absorbed it: the three earliest software-house drafts to the software-house supplement, the lead, and the testing supplement; The decisions that stay human, and what would move them to the bootstrap supplement, whose ordering principle its selection argument became; and The Bitter Lesson does not require everything to live in weights to the lead, which carries its rebuttal as a section.
Authoring and publication conventions live in COLLECTION.md; ADR 057 records the design.
Complete file listing (generated at build time)