← All papers
First page of Reinforced Large Language Model is a formal theorem prover

Reinforced Large Language Model is a formal theorem prover

Zhiling Luo

cs.AI Feb 13, 2025 · v1
Trains a Qwen2.5-0.5B model to predict Lean 4 tactics from Lean-workbook data, with proof search via LeanDojo evaluated on miniF2F.
To take advantage of Large Language Model in theorem formalization and proof, we propose a reinforcement learning framework to iteratively optimize the pretrained LLM by rolling out next tactics and comparing them with the expected ones. The experiment results show that it helps to achieve a higher accuracy compared with directly fine-tuned LLM.

LLM-based Lean theorem provers are usually fine-tuned directly from general base models, which limits their performance. The work asks whether reinforcement learning can improve next-tactic prediction for formal proving in Lean 4.

Lean-workbook samples are augmented with GPT-4o-generated chain-of-thought explanations of the ground-truth next tactic. A two-phase pipeline first applies supervised fine-tuning (the adaption phase) to Qwen2.5-0.5B. It then runs GRPO reinforcement learning with accuracy and format rewards, comparing rolled-out tactics against the ground truth. At inference time, a modified LeanDojo handles interaction with Lean, and tree search is adapted from InternLM.

Figure 1: The framework includes offline data preparation, model training and online inference.

On a 30-problem miniF2F subset, the RL model reaches 43% accuracy versus 36% for the SFT-only model. On the training set, accuracy rises from 49% to 61%.

Figure 3: Train process records of RL phase
Model#Positive on miniF2FAcc on miniF2FAcc on Trainset
Lean-qwen0.5B-sft1136%49%
Lean-qwen0.5B-rl1343%61%
Accuracy of SFT vs RL models on miniF2F (30 samples) and trainset