← All papers
First page of MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

Jiaxin Yuan, Connor Martinez Lockhart, Xiaoyu Liu, Jiaqi Wang, Chenghao Deng, Xiayimei Han, Vlassis Mastrantonis, Dmitrii Gudin, Shaopeng Zhu, Abdirisak Mohamed, Bilal Aytekin, Jiewen Lang, Zezheng Song, Furong Huang

cs.CL Aug 26, 2026 · v2 cs.AI cs.LO
Introduces MathAdv, a diagnostic benchmark of 298 problems formalized in Lean 4 across 13 domains, evaluating theorem-proving LLMs.
Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations. We introduce MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate- and graduate-level mathematics. Alongside Lean 4 theorem proving, MathAdv provides up to three auxiliary tasks: multiple-choice questions that probe mathematical knowledge, fill-in-the-blank problems that isolate informal reasoning, and expert-crafted transformations that test robustness to problem presentation. Our evaluation of contemporary theorem provers yields four findings: formalization remains a major bottleneck; performance varies substantially across mathematical domains; natural-language guidance helps general-purpose LLMs but can hinder proof-specialized models; and mathematically equivalent reformulations expose substantial robustness limitations. Together, these results show how component-wise evaluation can reveal model capabilities and failure modes that aggregate theorem-proving accuracy obscures. The dataset and evaluation scripts are available at https://github.com/margotyjx/MathAdv.git.

Existing formal mathematics benchmarks report only aggregate proof accuracy, cover a narrow range of domains (mostly algebra and number theory), and rarely test robustness to equivalent problem reformulations. This obscures whether failures stem from knowledge gaps, reasoning errors, or formalization difficulty.

MathAdv is a diagnostic benchmark of 321 problems (298 formalized in Lean 4) spanning 13 undergraduate- and graduate-level domains, built through a human-in-the-loop autoformalization pipeline. Each problem includes up to three auxiliary tasks: multiple-choice questions probing mathematical knowledge, fill-in-the-blank/direct-answer problems isolating informal reasoning, and expert-crafted transformations testing robustness. Contemporary proof-step and whole-proof provers plus general-purpose LLMs are evaluated on end-to-end Lean 4 proof construction alongside the auxiliary tasks.

Figure 3: Human-in-the-loop autoformalization process for MathAdv.

Formalization remains the major bottleneck (best model reaches 21.9% Lean 4 accuracy); performance varies substantially across domains; natural-language guidance helps general-purpose LLMs but can hinder proof-specialized models; and equivalent reformulations expose substantial robustness limitations.

Figure 5: Performance on Lean 4 formal proving with and without hints from MC questions.
ModelAccuracy (%)
DeepSeek-Prover-V1.5 RL + RMaxTS16.56
Goedel-Prover-V221.88
GPT-5.410.62
DeepSeek-R110.94
DeepSeek-V3.25.31
Lean 4 theorem-proving accuracy of selected models on MathAdv