FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games

Type: kb/sources/types/snapshot.md

Author: Leonardo Bertolazzi, Katya Tentori, Raffaella Bernardi Source: https://arxiv.org/abs/2606.04751 Date: June 3, 2026

arXiv ID: 2606.04751

Subject Category: Artificial Intelligence (cs.AI)

Abstract

This paper introduces FALSIFYBENCH, an evaluation framework designed to assess hypothesis-driven reasoning in large language models. The framework draws inspiration from the classic Wason 2-4-6 task, where agents discover hidden properties by proposing examples and receiving iterative feedback.

Key findings from evaluating 12 LLMs include:

  • Reasoning models demonstrate stronger scientific reasoning capabilities than instruction-tuned models, though none approach optimal performance
  • "The primary driver of success is the capacity for negative testing" — models actively seeking to falsify hypotheses outperform those seeking confirmation
  • Turn-level analysis reveals identifiable failure patterns in how models navigate hypothesis spaces

The task captures essential scientific reasoning components: hypothesis generation, evidence gathering, and belief revision based on both confirming and disconfirming evidence.

Access: PDF | HTML | TeX Source

License: CC BY 4.0