Ingest: FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games
Type: kb/sources/types/ingest-report.md
Source: falsifybench-inductive-reasoning-rule-discovery-games.md Captured: 2026-07-26 From: https://arxiv.org/abs/2606.04751
Classification
Genre: scientific-paper -- an arXiv (cs.AI) benchmark paper introducing an evaluation framework and reporting results across 12 models; the capture is abstract-level, so the genre is read from the artifact's form rather than from inspected methods. Domains: scientific-discovery, evaluation, reasoning, learning-theory Author: Leonardo Bertolazzi, Katya Tentori, Raffaella Bernardi (University of Trento). Bernardi is an established computational-semantics researcher and Tentori works on probabilistic reasoning and confirmation in cognitive psychology; the pairing is the right one for porting a Wason paradigm to LLMs. Independent academic group, no vendor stake in the result.
Summary
FALSIFYBENCH adapts the Wason 2-4-6 rule-discovery game into an interactive benchmark: an agent must identify a hidden semantic property by repeatedly proposing example instances and reading back whether each satisfies the rule. The paradigm exercises hypothesis generation, evidence gathering, and belief revision under both confirming and disconfirming feedback, in a closed loop where the agent chooses its own probes. Across 12 models spanning families and scales, reasoning models outscore instruction-tuned models but none approach optimal play. The paper's headline result is about how rather than how well: the primary driver of success is the capacity for negative testing — models that actively construct probes intended to falsify their current hypothesis consistently beat models that mostly propose confirming instances. A turn-level analysis, which the authors present as neglected in prior work, ties failure to identifiable patterns in how models traverse the hypothesis space rather than to a single aggregate score. Worth reading in full for anyone who needs the operationalization of "negative testing" or the failure taxonomy; the abstract alone already carries the load-bearing claim.
Connections Found
The KB has a live casebook here, and this source's role in it is process-level evidence for the falsification premise the discovery cluster currently asserts without external grounding. Three of its landing points are load-bearing rather than decorative.
Its strongest role is as evidence for first-principles reasoning selects for explanatory-reach over adaptive fit, whose third negative test ("can it be criticized?") and its operationalization in mechanistic constraints make Popperian KB recommendations actionable both rest on an untested premise about model behaviour: that criticism must be structurally forced because models will not seek disconfirmation ambiently. This paper is the first external measurement bearing on that premise, and it cuts both ways — falsification-seeking is confirmed as the discriminating behaviour, but the reasoning-model result shows the ambient capacity is real and varies by model rather than being uniformly absent. The Popperian note carries no external sources at all today, so this is its first empirical leg.
Second, it is a worked instance for known-target discovery benchmarks show reachability, not discovery closure, which currently runs on a single case (GIANTS, the backcast construction). FALSIFYBENCH is the authored-hidden-target construction the note distinguishes but does not instantiate — and it is a partial counter-case to the note's own framing, because the scored quantity is the agent's test-selection policy, not its recovery of the planted rule. Planting the target here buys a measurable process signal rather than converting discovery into target reconstruction.
Third, it measures the transition discovery lifecycle posits between consequence derivation and test — stating what would count against a conjecture, then going and looking. That phase boundary has been justified by Peirce and PDSA analogues; this is the first LLM-side data on whether the step actually happens in a closed loop.
Among sources it pairs most usefully with DiscoverPhysics — same closed experimentation loop, concrete simulated worlds instead of an abstract rule, same "best agents well short of optimal" shape — and with An Enigma of Artificial Reason, which finds the same confirmation-bias family at the evaluation locus where this paper finds it at the generation locus.
Extractable Value
- Falsification-seeking is the discriminating behaviour, not just a rhetorical virtue -- the KB's explanatory-reach machinery (the four-part negative test, falsifier blocks, the reach-assessment criterion) has argued this from first principles with no external support. A measured, cross-model result that negative testing is the primary driver of rule-discovery success upgrades that from asserted design taste to a claim with a behavioural correlate. [quick-win]
- A known-target benchmark can score the search policy instead of the answer -- this is the highest-reach item and the KB has not stated it. Where no outcome oracle for discovery exists, planting a target makes the agent's test selection measurable even though the final rule recovery is the trivially-known part. That partially escapes the "the benchmark already knows what counts as success" critique and generalizes past this paper: it is a construction rule for building oracles in oracle-poor domains, with FALSIFYBENCH and DiscoverPhysics as two worked cases. [deep-dive]
- Turn-level failure patterns beat aggregate scores for diagnosing hypothesis-space navigation -- the authors flag turn-level analysis as neglected in prior work. For a KB that cares about where a reasoning loop breaks rather than whether it passed, this is a reusable evaluation method: instrument the trajectory, taxonomize the transitions, and treat the aggregate as a summary of the trajectory rather than the measurement. [experiment]
- Ambient criticism capacity is model-dependent and non-zero -- "reasoning models are generally stronger scientific reasoners than instruction-tuned models, although no model comes close to optimal" bounds the structural-scaffolding argument from both sides. Structure is still needed (nobody is near optimal), but the premise that models never self-criticize without scaffolding is too strong, and which model runs a review gate is partly an empirical selection question. This bears directly on reach-assessment, which names the prose route to reach judgment as an open problem. [quick-win]
- A second Wason paradigm for the human-to-LLM transfer question -- the KB's transfer-boundary reasoning in human writing structures transfer to LLMs because failure modes overlap leans on the Wason selection task via Lampinen. The 2-4-6 rule-discovery task is the other paradigm in the family, and confirmation bias is the classic human failure on it, so the per-convention transfer question can now be asked on two tasks rather than one. Weakened by the capture reporting no human baseline. [just-a-reference]
- "Negative testing" as retrieval vocabulary -- a compact, greppable name for the behaviour the KB has been circling with "criticizability", "falsifier block", and "what would defeat this claim". Useful for discussion and search even where nothing else is imported. [just-a-reference]
Limitations (our opinion)
Editorial judgment, and constrained by a thin capture: the snapshot is abstract-level, with no methods, model list, numbers, or human baseline. Everything below is a caution about the claim as stated, not a finding about the paper's actual internals.
The load-bearing worry is that the headline result is at risk of being partly definitional. If "negative testing" is operationalized as proposing instances that fall outside the current hypothesis, then on a 2-4-6-style task where the hidden rule is characteristically broader than the natural first guess, probes outside the hypothesis are also the only probes that carry information. A model that tests outside its hypothesis would then score better because it gathered more evidence per turn, not because it holds a falsificationist disposition. That simpler account — negative testing as an information-gain proxy rather than an epistemic virtue — predicts the same correlation, and the abstract does not distinguish them. Whether the finding is hard to vary depends on details the capture does not carry: whether the authors controlled for probe informativeness, and whether the result survives rules where confirming probes are equally informative. Anyone citing this as support for the KB's negative test should read the operationalization before leaning on it.
Second, what was not tested, as far as the capture shows: a single task family. The 2-4-6 paradigm has a known quirk — the canonical hidden rule is deliberately more general than the seed example invites — and results on it have historically been sensitive to that framing. Twelve models across families is decent coverage, but one paradigm is not, and generalizing from "LLMs under-use negative testing on abstract semantic rule games" to "LLMs under-criticize their own claims in prose" is a jump the paper does not license. The KB's own systematic prompt variation serves verification and diagnosis, not explanatory-reach testing makes the parallel point about what varying an instance does and does not establish.
Third, no human baseline appears in the capture. Humans fail the 2-4-6 task badly and famously; without a baseline, "no model comes close to optimal" is a comparison to an optimal-play ceiling, not evidence that models are worse hypothesis-testers than people. Do not read it as the latter, and do not write that the paper compares LLM confirmation bias against the human literature — the capture does not support that claim.
Recommended Next Action
Update known-target discovery benchmarks show reachability, not discovery closure: add FALSIFYBENCH to the "Two target constructions" section as the reinvention-construction worked case beside GIANTS, and — the substantive edit — address the partial counter-case it raises, that a planted target can make the search policy scorable rather than only the reconstruction. That single revision lands the source's highest-reach contribution and gives the note the second instance it currently lacks; the reverse evidence edges from the explanatory-reach and Popperian notes can follow once the operationalization has been read in the full paper.
- known-target discovery benchmarks show reachability, not discovery closure — abstracted-from: FALSIFYBENCH is the authored-hidden-target construction the note distinguishes, and a partial counter-case because it scores test selection rather than target recovery
- first-principles reasoning selects for explanatory-reach over adaptive fit — evidence: the four-part negative test's criticizability criterion gets its first behavioural correlate
- mechanistic constraints make Popperian KB recommendations actionable — evidence: bears on the note's premise that criticism must be structural because it will not happen ambiently
- discovery lifecycle — evidence: first LLM-side measurement of the consequence-derivation to test transition
- reach-assessment — evidence: the prose route to reach judgment varies measurably across models
- conjecture is seeing the particular as an instance of the general — evidence: turn-level data on how the posit-and-recognize loop breaks
- DiscoverPhysics — compares-with: the same closed experimentation loop in simulated worlds rather than an abstract rule game
- An Enigma of Artificial Reason — compares-with: confirmation bias measured when models evaluate supplied reasoning rather than generate their own probes
- GIANTS — compares-with: the backcast construction of a known-target discovery benchmark
- Language Models, Like Humans, Show Content Effects on Reasoning Tasks — compares-with: the KB's other Wason-family source, on the selection task