EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
Mukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu, Shansan Gong, Qi Liu, Haitao Mi, Dong Yu
cs.CL
Sep 16, 2025 · v1
cs.AI
TL;DR
Trains LLM provers that generate whole Lean 4 proofs and evaluates them with the Lean checker on miniF2F and ProofNet.
Abstract
Large Language Models (LLMs) have recently advanced the field of Automated Theorem Proving (ATP), attaining substantial performance gains through widely adopted test-time scaling strategies, notably reflective Chain-of-Thought (CoT) reasoning and increased sampling passes. However, they both introduce significant computational overhead for inference. Moreover, existing cost analyses typically regulate only the number of sampling passes, while neglecting the substantial disparities in sampling costs introduced by different scaling strategies. In this paper, we systematically compare the efficiency of different test-time scaling strategies for ATP models and demonstrate the inefficiency of the current state-of-the-art (SOTA) open-source approaches. We then investigate approaches to significantly reduce token usage and sample passes while maintaining the original performance. Specifically, we propose two complementary methods that can be integrated into a unified EconRL pipeline for amplified benefits: (1) a dynamic Chain-of-Thought (CoT) switching mechanism designed to mitigate unnecessary token consumption, and (2) Diverse parallel-scaled reinforcement learning (RL) with trainable prefixes to enhance pass rates under constrained sampling passes. Experiments on miniF2F and ProofNet demonstrate that our EconProver achieves comparable performance to baseline methods with only 12% of the computational cost. This work provides actionable insights for deploying lightweight ATP models without sacrificing performance.
Problem
LLM-based automated theorem provers for Lean rely on test-time scaling, either long reflective chain-of-thought or many sampling passes. Both are computationally expensive, and existing cost analyses count only sampling passes, ignoring differences in token usage.
Approach
The authors quantify token-level sampling cost across scaling strategies and propose EconRL, a pipeline with two parts. The first is dynamic CoT switching, trained with DPO, so the model uses extended reasoning only on harder problems. The second is diverse parallel-scaled RL with trainable prefixes (proof heads), which raises pass rates when the sampling budget is small. EconRL is applied to DeepSeek-Prover-V2-7B and Goedel-Prover-V2-8B.
Results
On miniF2F and ProofNet, the resulting EconProver models reach accuracy comparable to full CoT mode at about 12% of the computational cost. For example, EconProver-GD scores 84.0% on miniF2F versus 84.4% for CoT mode, at 3x rather than 25x the relative cost of non-CoT mode.
| Model / Setting | Cost | Acc (%) | Rel. Cost |
|---|
| DeepSeek-Prover-V2-7B CoT mode | 4.5k × 32 | 75.8 | 10× |
| EconProver-DS | 1.2k × 16 | 76.2 | 1.5× |
| Goedel-Prover-V2-8B CoT mode | 12.3k × 32 | 84.4 | 25× |
| EconProver-GD | 2.9k × 16 | 84.0 | 3× |
miniF2F-test accuracy and cost (tokens per pass × passes; relative cost vs. non-CoT mode)