FrontierMath Erdős
Reported AI solutions to open mathematical problems suffer from reporting bias, undisclosed costs, unclear human scaffolding, and contamination, so they do not amount to a systematic measure of AI research capability.
68 open Erdős problems from erdosproblems.com were selected for difficulty and interest and autoformalized as Lean 4 statements using Mathlib, then reviewed by an expert. AI agents work in sandboxed Docker containers, with one attempt per problem, a $300 budget and 72 hours, and must prove either the conjecture or its negation. Submissions are accepted only if the Lean FRO's Comparator confirms the identical statement is proved with permitted axioms through the Lean kernel.
Five models were evaluated. GPT-6 Astra (pre-release) resolved 2 of 68 conjectures (#74 disproved, #126 proved), a 3% score, and all other models scored 0%. Additional non-systematic attempts also resolved #1, #548, and #571 at higher cost.
| Model | Score |
|---|---|
| GPT-6 Astra (pre-release) | 3% |
| GPT-5.6 Sol | 0% |
| GPT-5.5 | 0% |
| Claude Fable 5.1 | 0% |
| Claude Fable 5 | 0% |
