← All papers
First page of Autoformalization in the Era of Large Language Models: A Survey

Autoformalization in the Era of Large Language Models: A Survey

Ke Weng, Lun Du, Sirui Li, Wangyue Lu, Haozhe Sun, Hengyu Liu, Tiancheng Zhang

cs.AI May 29, 2025 · v2
Surveys LLM-based autoformalization and catalogs Lean 4 and Mathlib-derived datasets and benchmarks (LeanDojo, Lean-Workbook, miniCTX, LeanEuclid) alongside Isabelle and Coq work.
Autoformalization, the process of transforming informal mathematical propositions into verifiable formal representations, is a foundational task in automated theorem proving, offering a new perspective on the use of mathematics in both theoretical and applied domains. Driven by the rapid progress in artificial intelligence, particularly large language models (LLMs), this field has witnessed substantial growth, bringing both new opportunities and unique challenges. In this survey, we provide a comprehensive overview of recent advances in autoformalization from both mathematical and LLM-centric perspectives. We examine how autoformalization is applied across various mathematical domains and levels of difficulty, and analyze the end-to-end workflow from data preprocessing to model design and evaluation. We further explore the emerging role of autoformalization in enhancing the verifiability of LLM-generated outputs, highlighting its potential to improve both the trustworthiness and reasoning capabilities of LLMs. Finally, we summarize key open-source models and datasets supporting current research, and discuss open challenges and promising future directions for the field.

Translating informal mathematics into machine-checkable formal statements and proofs is labor-intensive. Recent LLM progress has led to rapid but fragmented work on autoformalization.

The survey organizes autoformalization research along three axes: mathematical domain, problem difficulty, and level of abstraction. It walks through the end-to-end workflow: data preprocessing, model design, post-processing and evaluation. It also covers the use of autoformalization to verify LLM outputs, including generated code. It compiles tables of open-source models and datasets, many of them Lean 4 or Mathlib-based, with notes on which informal and formal components each dataset contains.

It provides a categorized overview of methods and corpora such as MiniF2F, ProofNet, PutnamBench, LeanDojo, Lean-Workbook, FormL4 and Herald. It identifies data scarcity, generalization to abstract theories, and interactive hybrid systems as open challenges.

DatasetSizeDescription
MiniF2F488Olympiad and high-school/undergraduate statements
LeanEuclid173Euclidean geometry problems formalized in Lean
FormL417,137Informalized theorems from Mathlib 4
PutnamBench657Putnam competition problems
LeanDojo122,517Theorems, proofs, tactics, premises from mathlib4
Selected datasets and benchmarks surveyed