Measuring autonomy well enough to see it improve is an open problem
Type: kb/types/note.md · Tags: foundations, self-improving-systems
A shovel does none of the digging's judgment — every stroke is a human motor decision, the tool only transmits force. A hand-operated excavator still has a human deciding every bucket position; it amplifies force but not judgment. A computer-guided excavator that grades to a digital terrain model now has the machine choosing depth-at-this-point instead of the operator's hand choosing it. Somewhere in that progression "the machine now does part of the work" becomes true, but there is no principled place to draw the line and say "40% machine." Decision content is continuous; any percentage cutoff is arbitrary. The same problem applies to any claim that a system is "60% autonomous."
Actor allocation already avoids this trap for a single system by refusing the scalar: it reports who performs each named function of the improvement pathway — search, evaluation, retention, where the pathway is proposal-selection — against the same declared boundary membership is read against. That profile locates one system at one time but does not answer two comparative questions:
- Is this system becoming more autonomous over time? Tracking a profile across releases is tractable only after fixing its comparison grain. At a coarse functional grain, Commonplace's change-candidate survey treats
commonplace-freshness-statusas a new noticing channel alongside pre-existing mechanical checks,kb/log.md, and connect reports. That supports a change in mechanism and actor allocation inside an existing noticing/search function, not the addition of a new function. A coarse profile therefore remains commensurable but can hide meaningful mechanism-level reallocations; a finer profile records those channels but must explain how newly introduced coordinates compare with earlier states. A partial escape exists for this intra-system question: the boundary-contraction test asks whether a pathway completes with the human decisions withheld, indexed by scope and horizon — an existence claim about a run, never a comparison of function lists. Where contraction results differ, it orders two states of one system; it contributes nothing to the cross-system question. - Is one system more autonomous than another? A bare comparison needs both systems' profiles indexed by the same functions. Two architectural proposal-selection loops that both decompose into search, evaluation, and retention can be compared directly, in principle — the Gödel machine and Commonplace's own proposal-selection pathways are this kind of case. The Homeostat sits at a coarser floor: it can be reconstructed as variation, viability pressure, and retention, but not as the same architectural subtype because generation and rejection are not separated by an evaluator. But two systems need not decompose their work the same way at all: one KB tool might channel candidate formulation through typed collections with per-collection contracts, as Commonplace does; another might get an analogous quality effect through a single style guide and strict human review, with no separable "collection contract" step to count at all. Counting named functions and comparing the counts would be comparing units that do not correspond.
Neither difficulty is solved by adding detail alone. Two profiles need a commensurable decomposition before their per-function readings can be compared. Across time, a declared stable coarse grain may supply one at the cost of hiding finer reallocations. Across systems with different decompositions, no comparable grain has yet been established. Without such a grain, “is this getting more autonomous?” is not yet well-posed.
A count would also erase a real difference in stakes even where decompositions do match. Search and evaluation fail asymmetrically: a bad candidate that search produces unattended still meets evaluation and is rejected at the cost of wasted effort, while a bad acceptance that evaluation makes unattended becomes operative and nothing downstream catches it, since only the last filter's errors survive. Handing search to an agent is therefore comparatively cheap even without a strong local check; handing evaluation over is exactly where warrant, bounded by oracle domain, is the question that matters. A bare count of "how many functions run unattended" would treat these as interchangeable when they are not.
Open Questions
- Whether a coarse, largely system-independent function list — search, evaluation, retention, from the proposal-selection loop — can serve as a common ontology that most systems' finer decompositions refine, making cross-system and across-time comparison possible at that coarser grain even when finer function lists diverge.
- Whether a rough, admittedly imprecise proxy — counting non-human-performed functions, weighted by both scope and the search/evaluation stakes asymmetry above — is worth adopting despite lacking a principled basis, the way composite proxy scores are tolerated elsewhere for KB curation (notes need quality scores to scale curation).
Relevant Notes:
- Methodological and computational closure track different changes — grounds: actor allocation, the per-function, non-scalar autonomy profile this note takes as its starting point and finds insufficient for comparison
- A proposal-selection improvement loop requires search, evaluation, and operative retention — grounds: the function list that happens to be shared across the systems the grounding note already compares
- Self-improving system — grounds: the declared-boundary relativity the allocation profile inherits from membership
- Warranted autonomy is bounded by oracle domain — contrasts: that note bounds how much evaluation autonomy can be trusted; this note is about whether autonomy can be measured or compared at all
- False-positive generation is filtered; false-positive acceptance becomes operative — grounds: the search/evaluation stakes asymmetry a bare function count would erase
- Where change candidates come from in Commonplace — evidenced-by:
commonplace-freshness-statusis a new noticing channel inside a broader existing search/noticing function, exposing the choice between stable coarse profiles and finer mechanism-sensitive ones - Computationally directed self-improvement is a fixed-boundary reallocation ending in contraction — extends: the horizon-indexed existence test that partially answers the across-time question without a commensurable decomposition
- Ingest: Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems — see-also: calls for mathematical frameworks to understand and foster scientific human–AI synergy; the captured authors' abstract supplies no commensurable function decomposition or scalar contribution-attribution method, a negative bounded to that abstract rather than the uncaptured full paper
- Increasing computational autonomy relocates human effort to the frontier instead of reducing it — extends: the outcome-per-effort partial order, with total effort counted across configuration, review, recovery, and repair
- Tool usefulness, computational autonomy, warrant, and system power are separate dimensions — contrasts: separates four progress dimensions without solving any measure; the measurement problem here is one dimension's