← All papers
First page of MIRB: Mathematical Information Retrieval Benchmark

MIRB: Mathematical Information Retrieval Benchmark

Haocheng Ju, Bin Dong

cs.IR May 21, 2025 · v1 cs.CL cs.LG
Benchmark includes Lean-based retrieval datasets: informalized Mathlib4 statement search and LeanDojo premise retrieval from Lean 4 proof states.
Mathematical Information Retrieval (MIR) is the task of retrieving information from mathematical documents and plays a key role in various applications, including theorem search in mathematical libraries, answer retrieval on math forums, and premise selection in automated theorem proving. However, a unified benchmark for evaluating these diverse retrieval tasks has been lacking. In this paper, we introduce MIRB (Mathematical Information Retrieval Benchmark) to assess the MIR capabilities of retrieval models. MIRB includes four tasks: semantic statement retrieval, question-answer retrieval, premise retrieval, and formula retrieval, spanning a total of 12 datasets. We evaluate 13 retrieval models on this benchmark and analyze the challenges inherent to MIR. We hope that MIRB provides a comprehensive framework for evaluating MIR systems and helps advance the development of more effective retrieval models tailored to the mathematical domain.

Mathematical information retrieval underlies theorem search in libraries such as Mathlib4, answer retrieval on math forums, and premise selection for theorem proving. No unified benchmark existed for evaluating these varied tasks together.

MIRB collects 12 datasets across four tasks: semantic statement retrieval, question-answer retrieval, premise retrieval, and formula retrieval. Lean-specific components are Informalized Mathlib4 Retrieval and the LeanDojo premise retrieval set, where queries are Lean 4 proof states and the corpus is Mathlib4 declarations. Isabelle (MAPL) and HOL Light (HolStep) premise sets are also included. Thirteen sparse, open-source dense, and proprietary embedding models are evaluated with nDCG@10.

Figure 1: Overview of tasks and datasets in MIRB.

Models perform moderately on semantic retrieval tasks but poorly on reasoning-based tasks. Formal premise retrieval is hardest; LeanDojo scores fall below 7 nDCG@10 for the smaller models shown. Adding cross-encoder rerankers does not improve performance.

ModelInformalized Mathlib4LeanDojoAvg. (all 12)
BM2531.496.9132.23
gte-large-en-v1.538.053.7340.68
UAE-Large-V140.434.6438.27
bge-large-en-v1.541.995.4539.00
nDCG@10 for selected models (subset of columns)