Breadth-first inventory of curiosity questions
Purpose
The structural-coherence experiment studies one tractable failure: a locally acceptable unit does not trigger a wider question about its role. It should not define curiosity by fiat. Curiosity can also begin from an ambiguity, an unexpectedly regular pattern, a connection, an absent expected item, or a new possibility when nothing is visibly misplaced.
This file is a brainstorm inventory, not a taxonomy. Its rows are ordinary question-producing moves that may overlap, split, or disappear after comparison with real episodes. Do not coin durable names for them until a distinction changes an experiment, diagnosis, or design.
Elicitation requires maintained question-generation systems already owns the established argument that stored knowledge is inert unless a workflow asks activating questions, and that perspectives, checklists, adversarial prompts, incidents, disagreement, and audits can supply them. This workshop should not duplicate that claim. Its open problem is earlier and wider: which observations or relations can give rise to a valuable question, how an agent selects one without being handed the target, and whether the answer changes later work.
The inventory separates where a question comes from from what happens afterward. Different questions may enter the same downstream path:
possible question
-> significance and attention allocation
-> discriminating investigation
-> update, retention, deferral, or explicit rejection
-> possible reopening after new evidence
The existing reactive/prospective and paper-doubt/living-doubt distinctions concern parts of this path. They do not exhaust the sources of possible questions.
Working inventory
| Question-producing move | Working question | Local lead | Current standing |
|---|---|---|---|
| Encountered mismatch | What conflicts with the role, responsibility, prediction, or explanation expected here? | The local case corpus contains structural code and prose candidates. | Candidate cases exist, but no admitted online curiosity fixture yet. |
| Rule-boundary challenge | What would change if this otherwise useful rule did not apply here, and what suggests the case may be near its boundary? | ADR 042 records a worked counterexample to the exhaustive three-register claim. | Accepted historical episode; prospective agent origination remains untested. |
| Ambiguity separation | Which materially different interpretations fit the evidence, and what observation would distinguish them? | Changing requirements conflate genuine change with disambiguation failure shows the cost of silently selecting one reading. | Theory and retrospective examples exist; no question-origination trace has been isolated. |
| Impasse or confusion | What missing fact, distinction, operation, or representation prevents progress here? | The cognitive-architecture synthesis distinguishes impasse, confusion, prediction violation, repeated deliberation, and dissatisfaction as signals of different missing capacities. | Memory-first conjecture; no Commonplace comparison yet tests whether agents form the corresponding subproblem. |
| Capability or measurement limit | What could this mechanism, verifier, or evaluation distinguish even if it worked exactly as designed? | In the Decapod curiosity experiment, the maintainer asked what a compiled free-form constitution could achieve; the tested agent framings did not reliably reproduce that reasoning. | Strong local motivating episode and small negative capability lead; not a stable model result. |
| Alternative-mechanism comparison | What simpler or fundamentally different mechanism could produce the same result, and what value does the extra machinery buy? | The Decapod cost/benefit prompt reached the target finding in both of two trials. | Small positive prompt lead; replication and harder controls are missing. |
| Exception or scaffolding pressure | Why does this account need so many exceptions, categories, bridges, or repair rules? Would one missing distinction or mechanism explain the accumulation? | The linking-foundations episode turned repeated vocabulary repairs into a search for a generative theory; ADR 042 replaced an exhaustive taxonomy after a worked case exposed its cost. | Historical leads; no experiment yet separates useful complexity from missing-explanation pressure. |
| Missing expected item | If this account, inventory, or process were complete, what should be present but is not? Is the absence evidence, a capture gap, or a search failure? | The linking-foundations retrospective initially blamed missing capture, then showed that the premises were present but the historian's query and assembled conclusion were absent. | One unusually detailed corrective episode; the different causes of apparent absence must remain separate. |
| Unexplained regularity | Which repeated outcomes seem to share a mechanism even though none is individually a failure? | Short composable notes maximize combinatorial discovery cites improvement-log syntheses found by co-loading independent cases. | Retrospective examples, not a controlled discovery comparison. |
| Connection and recombination | What follows only when two previously separate artifacts, mechanisms, or perspectives are considered together? | The cognitive-architecture synthesis proposes iterative mapping in which representations and searches change together. | Memory-first conjecture; source grounding and a comparative task are still required. |
| Abstraction from particulars | What are these apparently different cases instances of, and where does the proposed common mechanism stop? | The improvement log records cross-artifact ABSTRACTION and SYNTHESIS candidates, while the linking-foundations work contains a longer human-agent abstraction episode. |
Several historical episodes, but no baseline distinguishes discovery from retrospective reconstruction. |
| Prediction-led surprise | What should happen next if the current model is right, and what should a violation teach us? | The HTM scan proposes attaching surprise to explicit workflow-transition predictions. | Memory-first method conjecture with no local experiment. |
| Newly available possibility | What useful action, representation, or question became possible only because this artifact, tool, or understanding now exists? | No clean local episode has yet been identified. | Open brainstorm item; currently outside the defect-centered evidence set. |
| Perspective or disagreement | What becomes visible from another stakeholder, time horizon, objective, or competent interpretation, and why do the accounts diverge? | The maintained-question-system note proposes perspective assignments and cross-model disagreement as question generators. | Existing methodological claim; the workshop has not isolated whether disagreement creates new questions or merely supplies more candidates. |
| Hidden contribution | What premise, judgment, objective, or competence is being supplied silently by a maintainer, model, environment, or evaluator? | The linking-foundations retrospective separates broad agent generation from a few load-bearing human inputs: query formation, significance, adoption, and stopping. | One detailed division-of-labor episode; causal generality is unknown. |
| Incidental discovery | What did the investigation find that was not its target, and should that finding alter the agenda or enter durable memory? | The Decapod runs found proof self-attestation, a no-op internalizer, the coplayer reliability system, and other bonus findings. | Discovery generation is observed; comparative valuation, uptake, and retention were not tested. |
| Inquiry allocation | Which possible question deserves scarce investigation effort because its answer could change a decision or discriminate important alternatives? | The parallel-terraced scan proposes expected-information-gain allocation, an exploration quota, and explicit stopping. | Memory-first conjecture; the Decapod runs only indirectly show breadth/precision tradeoffs. |
| Reopening | Which rejected, deferred, settled, or weak question should be reconsidered after new evidence changes its promise? | The maintenance-curiosity pilot freezes semantically old notes whose surrounding machinery may or may not have changed, alongside stable controls. | Runnable pilot, no result yet; the cognitive-architecture scan supplies only a memory-first policy conjecture. |
What the present workshop covers—and misses
The structural-misplacement track directly exercises encountered mismatch and rule-boundary challenge. Its stage-separated experiments also test significance, question persistence, candidate generation, selection, and uptake. That remains useful because the target can be made inspectable and controlled.
It does not yet test several materially different possibilities:
- a question generated by two live interpretations rather than a detected defect;
- an explanation sought for a positive regularity rather than a failure;
- a new inference produced by combining distant artifacts;
- an opportunity noticed because the available action space changed;
- an incidental finding competing with the commissioned target for attention;
- or an old question becoming valuable again after later evidence.
The original Decapod experiment also exposes a distinction the structural track currently underweights. “Does this mechanism work?”, “is this machinery worth its cost?”, and “what could this mechanism ever establish?” can inspect the same artifact while creating different rejection surfaces. The last question challenges the power of the mechanism or oracle itself; it is not merely another boundary case.
Question sources are not elicitation mechanisms
The inventory above describes what makes a question available. A workflow still needs some way to surface it. Existing Commonplace material suggests several mechanisms with different dependence on prior expertise:
- a direct targeted probe activates a question already known to the user;
- a perspective assignment or adversarial scenario supplies a reusable direction of search;
- a checklist carries previously learned questions into later cases;
- incident analysis turns an experienced failure into a question family;
- cross-model or cross-reviewer disagreement exposes a choice that one account alone can hide;
- periodic or randomly targeted audits commission inquiry without waiting for ambient human noticing;
- co-loading and connection mining place independent particulars together so a regularity can become visible;
- prediction, counterfactual simulation, and alternative generation create observations that ordinary execution never produces.
These mechanisms are not interchangeable evidence of autonomous curiosity. A direct probe demonstrates reachability; a checklist demonstrates reuse; an audit automates commissioning; disagreement creates a comparison surface; co-loading changes the available representation. For each episode, record both the source of the question and the mechanism that surfaced it.
Cross-cutting control questions
These questions apply after any source above produces a candidate inquiry:
- What material decision, explanation, or artifact change could the answer affect?
- What observable evidence made the question plausible rather than merely imaginable?
- What result would discriminate the live alternatives?
- What is the opportunity cost of investigating it now?
- Did the result change the artifact, plan, belief, uncertainty, or future search policy?
- If the inquiry stops, was it rejected, deferred for cost, or simply not reached?
- What later evidence would justify reopening it?
These controls prevent breadth from rewarding ornamental question lists. They also prevent the opposite mistake: treating only questions with an immediate defect and deterministic test as curiosity.
Candidate experiment clusters, not commitments
The inventory currently suggests three different clusters worth comparing before a broad theory is attempted:
- Defect- and limit-seeking: encountered mismatch, rule-boundary challenge, ambiguity separation, missing expected items, and capability or measurement limits. The structural-coherence fixtures and Decapod episode are the strongest current substrates.
- Generative discovery: unexplained regularity, connection and recombination, abstraction, prediction-led surprise, and newly available possibilities. The KB has retrospective examples but almost no prospective baselines.
- Inquiry control over time: hidden contributions, incidental discovery, allocation, persistence, stopping, retention, and reopening. The linking retrospective and Decapod bonus findings expose the problem, while the cognitive-architecture scans offer ungrounded process conjectures.
These clusters are only a way to choose contrasting experiments. They should not become a curiosity type system. The next selection should prefer phenomena that differ in what triggers the question, have a recoverable local episode, and admit an observable downstream consequence.
Immediate gaps
- No local episode clearly records opportunity-seeking curiosity: an agent noticed what newly became possible without first being shown a problem.
- No experiment tests whether an agent can discover a regularity or connection prospectively rather than reconstruct a supplied target.
- No local comparison tests perspective assignment or disagreement as a source of genuinely new questions rather than additional critiques.
- No trace follows a valuable incidental finding through prioritization, investigation, retention, and later reuse.
- No completed local case yet tests reopening after evidence changes the value of an earlier question; the maintenance-curiosity pilot now supplies the first frozen cohort.
- The structural track has a detailed stage model; the generative and longitudinal tracks do not yet have comparable evidence and should not inherit that model without testing.
Next use
Before expanding the main experiment, select two or three rows with different question sources and prepare one-page episode records containing:
- the situation before the question was asked;
- who or what originated it;
- the evidence that made it worth asking;
- the investigation and possible rejection surface;
- the update or failure to update; and
- the strongest rival account of the episode.
The likely first records are the Decapod capability-limit/cost comparison, the linking-foundations query-formation and valuation episode, and the ADR-042 counterexample. Their differences are more informative than adding more structural-placement examples immediately.