← All papers
First page of FrontierMath Erdős

FrontierMath Erdős

Tom Adamczewski, Thomas F. Bloom

cs.CL Sep 6, 2026 · v1 cs.AI
Benchmark of 68 open Erdős problems stated in Lean 4 with Mathlib; AI agents must produce Lean proofs, which the Comparator tool checks.
We introduce FrontierMath Erdős (FME), a benchmark of 68 Erdős problems that are open as of August 2026. To solve a task in FME, AI systems must resolve (prove or disprove) one of the 68 conjectures in the proof assistant Lean. Our 68 problems were selected by the second author among 652 open problems on erdosproblems.com for their mathematical interest and difficulty. AIs have recently resolved several open problems in mathematics, but these demonstrations fall short of a systematic study of AI capabilities. FME evaluates every AI model on the same fixed problems, autonomously and under the same budget. We evaluated five AIs with a budget of \$300 per problem. One (GPT-6 Astra) scored 3%, and all others scored 0%.

Reported AI solutions to open mathematical problems suffer from reporting bias, undisclosed costs, unclear human scaffolding, and contamination, so they do not amount to a systematic measure of AI research capability.

68 open Erdős problems from erdosproblems.com were selected for difficulty and interest and autoformalized as Lean 4 statements using Mathlib, then reviewed by an expert. AI agents work in sandboxed Docker containers, with one attempt per problem, a $300 budget and 72 hours, and must prove either the conjecture or its negation. Submissions are accepted only if the Lean FRO's Comparator confirms the identical statement is proved with permitted axioms through the Lean kernel.

Five models were evaluated. GPT-6 Astra (pre-release) resolved 2 of 68 conjectures (#74 disproved, #126 proved), a 3% score, and all other models scored 0%. Additional non-systematic attempts also resolved #1, #548, and #571 at higher cost.

ModelScore
GPT-6 Astra (pre-release)3%
GPT-5.6 Sol0%
GPT-5.50%
Claude Fable 5.10%
Claude Fable 50%
Benchmark scores (one attempt, $300, 72 h per conjecture)