Ingest: Safe superintelligence
Type: kb/sources/types/ingest-report.md
Classification
The chapter reviews formal intelligence measures and the superintelligence/safety canon, then advances a philosophical critique of definitions, benchmarks, and speculative claims. Author: Richard Heimann is a secondary interpreter of Legg, Hutter, Good, Vinge, Bostrom, Yudkowsky, Turing, and philosophical traditions. The genealogy is useful orientation; it is not empirical evidence for hard takeoff or particular risk probabilities.
Summary
The final chapter starts with Shane Legg's performance-based definition of intelligence, universal intelligence measure, and AIXI, an idealized reward-maximizing agent over computable environments. It then traces the safety canon from intelligence explosion and singularity through superintelligence and “foom,” before criticizing definitions that masquerade as explanations and metrics that quietly inherit behaviorist or functionalist commitments. Its proposed alternative is procedural: name the game, state public criteria and limits, keep philosophical possibility separate from claims used for engineering or policy, and preserve oversight where objectives, reward channels, environments, or robustness remain uncertain.
Quotes
No source quotes have been retained yet.
Connections Found
The universal measure is a strong case for a proximate target being checked for achievement rather than warrant: formal score maximization does not prove that the score captures understanding, value, or safety. AIXI's exogenous reward and environment assumptions support warranted autonomy being bounded by oracle domain. Recursive-improvement rhetoric compares with self-improvement relative to a declared objective and computational direction as fixed-boundary reallocation: without a boundary, horizon, objective, and allocation trace, “foom” is not an operational improvement account.
Extractable Value
- Operationalization does not discharge ontological or normative warrant -- a universal score can rank agents under its formal assumptions while leaving “is this intelligence?” and “is this good?” as separate linking claims. [deep-dive]
- An oracle domain includes reward integrity -- AIXI assumes an exogenous reward channel; real autonomous systems may influence evaluators, observations, goals, or environments, shrinking the domain in which optimization is warranted. [quick-win]
- Recursive self-improvement needs a fixed frame -- compare the same boundary, objective, and horizon while tracking which search, evaluation, and retention decisions move to computation; otherwise category changes can be manufactured by redrawing the system. [quick-win]
- Public tests are provisional contracts, not explanations -- procedures and benchmarks make disagreement actionable, but their results must not be inflated into consciousness, understanding, or universal competence. [just-a-reference]
- Genealogy is not probability evidence -- historical influence from Good through Bostrom and Yudkowsky explains why a scenario is salient, not how likely it is or whether proposed controls work. [quick-win]
- Audit intelligence/safety claims by requirement chain -- name the formal metric, claimed capability, deployment objective, oracle domain, and untested link at each step. [experiment]
Limitations (our opinion)
The chapter correctly separates tests from essences but sometimes replaces one broad framing with another: “procedural, not metaphysical” does not tell us which procedures are sufficiently discriminating. Its safety history emphasizes canonical texts and citation influence, which can create an impression of cumulative evidence where there is mainly cumulative discourse. Idealized AIXI results omit computational limits and manipulable feedback; conversely, pointing out those omissions does not establish low real-world risk. The treatment of functionalism, behaviorism, consciousness, and meaning is compressed enough that labels occasionally do more explanatory work than arguments. No empirical evidence in the chapter warrants a hard-takeoff probability or validates a control proposal.
Recommended Next Action
Use Legg's universal intelligence measure as a worked case in the next revision of A proximate target is checked for achievement, not for warrant, but ground the example in Legg's dissertation and keep superintelligence probability claims out of that revision.