← All papers
First page of Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

Lan Zhang, Marco Valentino, Jordan Meadows, Andre Freitas

cs.CL Jun 12, 2025 · v2
Evaluates LLM-judge ensembles on Lean4 and Isabelle/HOL autoformalizations from miniF2F and ProofNet, with formal validity checked by the theorem prover.
Statement autoformalization plays a crucial role in formal mathematical reasoning by enabling the automatic translation of natural language statements into formal languages. While recent advances using large language models (LLMs) have shown promising capability of autoformalization, methods for automatically evaluating autoformalization remain underexplored. LLM-as-a-judge presents a promising approach for automating such evaluation, however, existing methods typically employ coarse-grained and generic evaluation criteria, which limit their effectiveness for advanced formal mathematical reasoning, where quality hinges on nuanced, multi-granular dimensions. In this work, we take a step toward addressing this gap by introducing a systematic, automatic method to evaluate autoformalization tasks. The proposed method is based on an epistemically and formally grounded ensemble (EFG) of LLM judges, defined on criteria encompassing logical preservation (LP), mathematical consistency (MC), formal quality (FQ), and formal validity (FV), resulting in a transparent assessment that accounts for different contributing factors. We validate the proposed framework to serve as a proxy for autoformalization assessment within the domain of formal mathematics. Overall, our experiments demonstrate that the EFG ensemble of LLM judges is a more suitable emerging proxy for evaluation than a coarse-grained model. These findings suggest that LLM-as-judges, especially when guided by a well-defined set of atomic properties, could offer a scalable, interpretable, and reliable support for evaluating formal mathematical reasoning.

Evaluating autoformalization, the translation of natural-language math statements into formal languages, is hard. Syntactic checks miss semantic errors, human review does not scale, and reference-based metrics depend on ground truths that may be wrong.

The authors propose an epistemically and formally grounded (EFG) ensemble of LLM judges. Each judge scores one atomic criterion: logical preservation, mathematical consistency, formal quality, or formal validity, with the last checked by the theorem prover. A linear model combines the criteria into an overall score. Human rankings and assessments of formalizations from miniF2F and ProofNet in Lean4 and Isabelle/HOL are used for validation.

Figure 1: Overview of the proposed framework for automatically evaluating autoformalization. Our method leverages ensembles of LLM judges based on epistemic criteria organized into a taxonomy, while also providing explanations for their judgments. A linear evaluation model aggregates the criteria into a single overall assessment to serve as a proxy for human evaluation.

Fine-grained EFG judges agree more closely with human rankings than coarse-grained judges. GPT-4.1-mini guided by operable atomic properties reaches about 72.7% pairwise ranking match with one human annotator on miniF2F, comparable to the 70% agreement between the two humans. Human review also found that many ground-truth formalizations in the benchmarks are flawed.

AFJudgeLPMCFQOA
GPT-4.1GPT-4.1-mini0.9240.9850.9570.891
Qwen-7BGPT-4.1-mini0.6270.8420.7580.703
GPT-4.1Qwen-Coder-7B0.7030.7570.3280.621
Qwen-7BQwen-Coder-7B0.4730.6160.1700.461
Overall-assessment (OA) agreement on miniF2F test set with different autoformalizers (AF) and judges (Lean4)