← All papers
First page of Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery

Syed Rifat Raiyan, Mohsinul Kabir, Hasan Mahmud, Md Kamrul Hasan, Sophia Ananiadou

cs.AI Jun 7, 2026 · v4 cs.CL cs.CV cs.LG
Survey that devotes one of its four axes to Lean 4 autoformalization, tactic prediction and proof search, with a Lean/Mathlib example.
Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem within NLP to one of the most consequential AI frontiers. This survey provides a unified account of the field's evolution, from early rule-based math word problem (MWP) solvers and template-driven geometry systems, through neural expression generation and LLM prompting, to contemporary reasoning models, multi-agent systems, neuro-symbolic theorem provers, and verified discovery workflows. We organize the landscape along four axes: (i) informal reasoning over text and diagrams, spanning MWP solving, multimodal geometry, and VLMs; (ii) formal reasoning in proof assistants, including autoformalization, tactic prediction, compiler-guided repair, and proof search; (iii) mathematical discovery, where systems propose constructions, improve bounds, or assist attacks on open problems; and (iv) the inference and training-time techniques, including CoT prompting, tool use, process reward models, and RLVR, that increasingly connect generation with verification. We catalog major benchmarks across grade-school arithmetic, competition mathematics, geometry, formal proving, multimodal and multilingual reasoning, and expert evaluation, and we examine benchmark saturation, contamination, reporting mismatches, and the distinction between pass@1, majority voting, and verifier-assisted pass@$k$. We critically assess failure modes: brittleness under perturbation, reward hacking, multimodal grounding failures, fragile formalization, and the energy cost of reasoning-scale inference. Drawing on recent perspectives from working mathematicians, we identify future directions centered on verified-discovery workflows, reasoning efficiency, and infrastructure to make AI-assisted formalization broadly usable. Companion materials: https://github.com/Starscream-11813/awesome-AI4Math.

AI for mathematical reasoning has expanded from rule-based word-problem solvers to reasoning models, formal provers and discovery systems. The field lacks an integrated account connecting these threads.

The survey organizes the literature along four axes. These are informal reasoning over text and diagrams, formal reasoning in proof assistants (mainly Lean 4: autoformalization, tactic prediction, compiler-guided repair, proof search), mathematical discovery, and inference/training techniques such as chain-of-thought, process reward models and RLVR. It catalogs benchmarks and analyzes saturation, contamination and differences in evaluation protocols.

It provides a taxonomy, dataset catalog, performance comparisons and a critical assessment of failure modes. It identifies future directions centered on verified-discovery workflows, reasoning efficiency, and infrastructure for AI-assisted formalization around Lean 4.