Introduces SA-Pass metric and ShadowBench, a Lean 4 autoformalization benchmark using shadow theorems verified via the Lean compiler and Mathlib.
Abstract
Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluation metrics can accept type-correct but misaligned statements or reject correct statements written in a different formulation. Inspired by Pass@$k$, we propose SA-Pass (*Semantic Alignment Pass*), which tests formal statements using auxiliary statements called *shadows* that characterize the intended statement. A generated statement receives full credit only when it compiles, implies each shadow (forward check), and is implied by their conjunction (backward check). We instantiate SA-Pass in ShadowBench, a Lean 4 full autoformalization benchmark of 178 postgraduate- to research-level problems spanning eight mathematical areas. Claude Code (Opus 4.8) with Numina-Lean-Agent reaches $61.8\%$ compile rate and $11.2\%$ SA-Pass. Across outputs generated by six agentic configurations, SA-Pass achieves $98.8\%$ binary agreement with expert judgments. An early version of ShadowBench served as the benchmark for Track 4 of the ICML 2026 AI4Math Challenge.
Problem
Autoformalization translates informal theorems into proof-assistant code, but current evaluation metrics accept type-correct but misaligned statements or reject correct statements in alternative formulations. Manual expert judgment is reliable but too costly to scale.
Approach
SA-Pass (Semantic Alignment Pass) tests a generated formal statement against auxiliary statements called shadows characterizing the intended theorem. Full credit requires the statement to compile, imply each shadow (forward check), and be implied by their conjunction (backward check), all verified in Lean 4 with Mathlib. This is instantiated in ShadowBench, a benchmark of 178 postgraduate- to research-level problems across eight mathematical areas, with LLM-assisted, compiler-verified checker construction.
Figure 3: Evaluation procedure for SA-Pass and SA-Pass soft on one ShadowBench problem.
Results
Claude Code (Opus 4.8) with Numina-Lean-Agent reaches 61.8% compile rate but only 11.2% SA-Pass, showing compilation overstates semantic alignment. SA-Pass achieves 98.8% binary agreement with expert judgments (F1 0.964, precision 1.000, recall 0.930), outperforming compile rate, BLEU, BEq+, and LLM-as-judge.
Metric
Precision
Recall
F1
Agreement
Compile
0.178
1.000
0.302
0.178
BEq+
0.400
0.186
0.254
0.806
SA-Pass soft
0.414
0.953
0.577
0.752
SA-Pass
1.000
0.930
0.964
0.988
Agreement of automatic metrics with expert judgment