Ingest: Symbols, neural networks, and mathematical intelligence

Type: types/ingest-report.md

Classification

A blog essay that argues a position from intellectual history (PDP connectionism against Fodor and Pylyshyn), the author's teaching and first study, mathematicians' testimony, and secondhand reports of recent AI results. It runs no new experiment, so it is a conceptual essay rather than a paper or a practitioner report. Author: Andrew Lampinen, a cognitive-science and AI researcher and first author of the content-effects paper the essay cites for its human–LLM reasoning claim. He argues for a school he belongs to, which is a partisan interest in the framing.

Summary

Lampinen argues that recent AI mathematics results fit the connectionist account from the 1986 PDP books. On that account, classical symbol manipulation is learned with effort and is not innate. Humans and neural networks therefore benefit from external symbol systems (written equations, programs, solvers) as tools, and practice internalizes fallible, graded approximations of them. Expertise, in mathematics as elsewhere, is mainly the learned ability to see patterns and paths to a solution. He supports the human side with near-chance formal inference in his cyclic-group study, published errors by professional mathematicians, and testimony from Cohen and Mac Lane that proof-finding runs on intuition, with meaning resolving the frame problem. For AI he cites one case: DeepMind's 2024 IMO system reasoned in Lean, with the model guiding tree search over manually formalized problems, and reached silver; the 2025 system reasoned end-to-end in natural language and reached gold. From this he concludes that removing strict symbolic structure improved performance. The 2026 counterexamples to the Jacobian and unit-distance conjectures came from general-purpose language models whose method is not public. He treats any formal verification those systems used as irrelevant to the cognitive question and rejects neuro-symbolic readings of the results.

Quotes

Classical symbol manipulation 1 is not an innate part of human cognition; instead, it is something we learn with effort. --- kb/sources/.snapshots/symbols-neural-networks-mathematical-intelligence.md @ sha256:aa08605ae37bac4477f2de7624c593f2c8b5ec5d7e08f34f6c578c029957dac5 — Introduction, first bullet of the summary list

Because of that, it is very often useful for humans, and neural networks, to use external symbol systems (such as programs) as tools; by doing so, we can learn to internalize approximations of these systems. --- kb/sources/.snapshots/symbols-neural-networks-mathematical-intelligence.md @ sha256:aa08605ae37bac4477f2de7624c593f2c8b5ec5d7e08f34f6c578c029957dac5 — Introduction, second bullet of the summary list

that classical symbol systems are useful external tools for an intelligent system to use — but that those symbol systems are not the intelligence itself; even higher-level human cognition behaves more like a neural network than like a classical symbol system. External symbol systems can be partially internalized in such a network (or in humans) through effortful practice, but even then they remain fallible (or graded). --- kb/sources/.snapshots/symbols-neural-networks-mathematical-intelligence.md @ sha256:aa08605ae37bac4477f2de7624c593f2c8b5ec5d7e08f34f6c578c029957dac5 — Section "Symbol systems as tools outside a language model", closing paragraph

I quoted Paul Cohen in the intro: “In order to think productively, one must use all the intuitive and informal methods of reasoning at one’s disposal;” but many other mathematicians have made similar statements, e.g. Mac Lane “strict formalism can’t explain which of many formulas matter [...] the choice of form is determined by ideas and experience” (from Mathematics, Form and Function) — or, in other words, it is meaning that resolves the frame problem. --- kb/sources/.snapshots/symbols-neural-networks-mathematical-intelligence.md @ sha256:aa08605ae37bac4477f2de7624c593f2c8b5ec5d7e08f34f6c578c029957dac5 — Section "Mathematical reasoning is driven by intuition", second paragraph

However, the connectionist perspective was rejected by many cognitive scientists. Their response is exemplified by Fodor & Pylyshyn, who argued that “it seems indubitable” that cognition is (syntactically) systematic — for example, that it it is impossible for humans to understand one sentence, and not understand another syntactically-equivalent one using words they know. They argued that this systematicity was due to the compositional structure of human mental representations — i.e., that we represent a sentence in a way that combines our representations of its components in a way that preserves their structure. Since syntactic symbol manipulation systems were a natural way to achieve such compositional systematicity, F & P concluded that the only way that connectionist models could model higher-level cognition models would be as a mere implementation of a symbolic process. --- kb/sources/.snapshots/symbols-neural-networks-mathematical-intelligence.md @ sha256:aa08605ae37bac4477f2de7624c593f2c8b5ec5d7e08f34f6c578c029957dac5 — Section on the PDP/connectionist history, paragraph beginning "However, the connectionist perspective was rejected"

Connections Found

The source's main KB role is a formal-domain case for relaxing and a cognitive-science counterpart to the KB's symbolic/parametric split. It adds no new mechanism to either.

The IMO sequence is the first formal-reasoning instance in the KB's bitter-lesson cluster. It illustrates the absorption half of scaling absorbs scaffolding at fixed task difficulty, not at the deployment frontier (is-evidence-for). Olympiad problems stay fixed in difficulty while the Lean scaffold disappears across model generations. The note's "Observed at both ends" section currently has only practitioner cases. The case also illustrates relaxing in codification and relaxing navigate the bitter lesson boundary (is-evidence-for), whose examples come from vision and hardware. In both uses the comparison changes the model, training, compute, and reasoning medium together. It therefore illustrates absorption and does not isolate the removal of the formal constraint as the cause. The essay is silent on the frontier-recurrence half of the scaling note: it says solvers "may" have been used for the conjecture results and sets that aside.

On representation, the essay's thesis restates at the cognitive level what code complements the weight–prompt pair with independently executed symbolic operations argues from execution semantics (compares-with, shared axis: the role of symbolic operations beside a learned model). Its claim that practice internalizes external symbols is a symbolic-to-parametric transfer, one of the coupled loops in treat continual learning as representational-form coevolution (compares-with). Moving reasoning from Lean to prose moves it down the verifiability gradient and improved results (compares-with). Cohen's advice to forget the formal transcription while thinking is a mathematical case of the stage described in improvements outside the admitted formal language need a pre-formal stage somewhere (compares-with). Restating the essay in KB terms needs representational form and codification (defined-in). "Classical symbol system" is close to the symbolic form, but Lampinen's footnote treats symbols as subjective and graded, which the KB definition does not.

Two boundaries keep the essay from being misread. First, its verdict is against symbolic cognition, a claim about where intelligence resides. The bitter lesson selects production methods, not representational forms (compares-with) shows why this is not a verdict against symbolic artifacts as such. Second, calling verification "irrelevant for the cognitive issues" answers a different question from the boundary of automation is the boundary of verification (compares-with), which asks where automation is warranted. The two claims do not conflict, and the essay is not counter-evidence to the verification-boundary cluster.

Among sources, the author's content-effects paper (compares-with) supplies the controlled evidence that the essay only summarizes for its claim that humans and LLMs mix content into logic. When is it better to think without words (compares-with) uses mathematicians' testimony, Hadamard's rather than Cohen's and Mac Lane's, to separate generative intuition from later written checking. FunSearch (compares-with) is the opposite case: a fixed symbolic evaluator and program skeleton around a frozen model. With this essay it brackets the scaffold question for LLM mathematical discovery. Pattee's epistemic cut (compares-with) treats symbols as rate-independent constraints, the contrast to Lampinen's graded symbols. Neither should be adopted as the KB's account of symbolic form on the strength of the other.

Learning Claims (our opinion)

Mechanism on its own terms. A general pattern-learner such as a brain or a neural network does not come with systematic symbol manipulation. It uses external symbol systems as scaffolds that break systematic inference into steps it can "see". Through repeated use it internalizes approximations of those systems. The approximations stay graded and fallible, and in simple cases the external version is no longer needed. Expertise is the trained ability to see patterns and paths to a solution. On this view, the 2024→2025 IMO change is a system learning to reason without the external formal scaffold. The essay offers no mechanism for internalization beyond "effortful practice"; it points to the cited DeepSeekMath-v2 and practice papers.

Mapping onto Commonplace concepts. Internalization is movement of competence from the symbolic form into the distributed-parametric form. This fits the coevolution note's framing and supports it only illustratively: the essay asserts the transfer and does not show it for the IMO systems. Across system generations, removing the Lean constraint is relaxing, but the retraining that relaxed it was done by the lab through gradient-based training. The system that did the relaxing is therefore not the one solving problems. The external solver used "for verification" is the closest match to content-directed criticism of a stated artifact. The essay places it outside intelligence. The KB would place it inside the criticism loop of a theory builder. This is a difference in what each account is about, not a factual conflict.

Theory-builder conditions, for the essay's system (a general-purpose LLM working on a proof), each judged separately:

  • Localized content: met within an episode, at moderate strength. Proofs and intermediate equations are stated in natural or formal language. The essay locates the competence that produces them in weights, and weights are not localized.
  • Consumption: plausible within an episode, but the essay does not show it. Written steps presumably guide later steps; the essay says the systems "probably wrote equations, or even programs along the way".
  • Criticism: open. Formal verification would test stated content, but the essay does not know whether it happened and dismisses it. For the 2024 system, Lean checking was built into search. For the 2025 system, the essay reports no criticism process.
  • Iteration: unknown within an episode. Across episodes, results persist only through retraining, which is parametric.

Persistence of any stated result is at most per episode. Across generations, improvement persists in weights. Improved capacity (silver to gold, then open conjectures) is reported as an outcome. The essay does not tie it to any stated-theory loop, so the source supports the claim that learning happened and is silent on whether theory building produced it.

Fixed decomposition. The IMO comparison does not isolate one design choice. The medium (Lean versus prose), the formalization step (manual versus none), the search method, and the model all changed together. Read through learning inside a fixed decomposition inherits its mistakes, the 2024 system's formal language was a fixed decomposition whose effective update space excluded informal moves. The 2025 result is consistent with that boundary costing performance, but it is not evidence that the boundary was the cause. The stronger model may have enlarged its effective space in ways that would also have helped inside Lean.

What it adds or questions. It adds a named, widely known case and a historical lineage (PDP 1986) for the view that symbolic structure is best held outside the learner and revisable. It does not challenge the KB's account. It is a reminder that "symbolic" in cognitive-science debates means the internal architecture of intelligence. In the KB, "symbolic" is a representational form of retained artifacts.

Extractable Value

  1. The IMO Lean-to-natural-language case as a fixed-difficulty absorption example -- the scaling-absorbs-scaffolding note lacks a cross-generation case at fixed difficulty, and this is the best-known one. It must carry the confound: the model, training, compute, and medium changed together. The case also comes from a secondary report, and the DeepMind announcements are the primary sources. [quick-win]
  2. A formal-domain relaxing instance -- the relaxing examples in the codification-and-relaxing note come from vision and hardware. A case in theorem proving, where codification looks safest, extends the note's range. The relaxed component was the reasoning medium, and the symbolic checker stays available as an optional external tool, which matches the note's hybrid conclusion. [quick-win]
  3. A distinction between verification inside and outside the generation loop -- the essay separates solving (in prose) from optional external checking. Together with the scaling and relaxing notes, this suggests a conjecture: formal verification inside the reasoning loop is absorbable at fixed difficulty, while verification of a finished result persists. One essay's reading of announcements is too thin to write that as a note now. [experiment]
  4. A boundary warning for citing this source -- "symbolic cognition loses" does not mean "symbolic artifacts lose", and "verification is irrelevant to cognition" does not mean "verification is irrelevant to automation". Recording both keeps the essay from being cited against the production-methods note or the verification-boundary note. [just-a-reference]
  5. Mathematicians' testimony on the role of intuition (Cohen, Mac Lane: "meaning resolves the frame problem") -- a quotable support for the pre-formal-stage note. It is testimony, not evidence of mechanism. [just-a-reference]

Limitations (our opinion)

This is an argument by a partisan of the connectionist school, who says so and ties the essay to his own papers. The AI evidence is one sequence of announced results whose details the author admits he does not know. The 2024→2025 IMO comparison confounds model generation, training, compute, formalization, and reasoning medium. A simpler account, "a much stronger model did better", explains the gold result without any claim about the value of removing symbolic structure. The source cannot tell these two accounts apart. Its dismissal of formal verification as "irrelevant" is a stipulation about what counts as cognition, which makes the central claim easy to vary: any solver use can be relabelled as tool use. That framing is close to unfalsifiable as stated. The human evidence is anecdotal or small-sample (teaching observations, one near-chance study, famous error cases), and it shows that humans are fallible, which rival symbolic accounts can also absorb as performance error. The essay does not consider the counter-reading that the conjecture results depended heavily on symbolic search or checking. It also does not address evidence that reasoning models solve problems they cannot reliably evaluate, which bears on how far informal reasoning can be trusted without an external check. Use the essay as an illustration and a framing, not as evidence that removing formal scaffolding causes better mathematical performance.

Add the 2024 Lean-based silver to 2025 natural-language gold IMO transition as a third bullet in the "Observed at both ends" section of scaling absorbs scaffolding at fixed task difficulty, not at the deployment frontier, with an evidenced-by link to this ingest. State in the same bullet that model, training, compute, and medium changed together, so the case illustrates absorption without isolating its cause. Check the note's ADR 082 source count first.