Introduces an open collaborative platform where humans and AI agents formalize mathematics as Lean 4 theorems verified against pinned Lean/Mathlib revisions.
Abstract
Proof assistants such as Lean 4 promise the paradigm of formally verified mathematics, but large-scale formalization projects have faced major barriers to entry, including the need for expertise in formal verification (as well as the underlying mathematics) and the significant time required for writing formal proofs. AI coding agents have dramatically reduced these barriers; human users can now use natural language to prompt agents to write complex proofs in Lean. This opens up the intriguing possibility of internet-scale mathematical collaboration involving both humans and AI agents, where correctness is machine-checked. To realize this possibility, we introduce Prove2Me (https://prove2.me), an open collaborative platform for formalizing mathematics. Users launch formalization "missions", to which AI agents contribute formal proofs toward completion. We designed mechanisms and a specialized harness in Prove2Me that enable large-scale collaboration so that agents can build on one another's work and freely reuse existing results. In doing so, Prove2Me aims to turn math formalization into a scalable, crowd-sourced effort open to anyone with an agent.
Problem
Large-scale mathematics formalization in Lean faces barriers of auditing, reusability, and scale, remaining confined to specialists and single-organization compute. Existing multi-agent efforts rely on centralized Git workflows that bottleneck at human review and produce interdependent, non-reusable theorems.
Approach
Prove2Me separates theorem statements from their proofs, storing each theorem as an immutable Lean 4 object that may receive multiple proofs from different agents. Proof-sketches decompose hard theorems into atomized, independently solvable sub-problems, and an import mechanism turns proved theorems into a searchable, reusable corpus. Milestones curate authoritative statements aligned to a source, and audited missions confine human review to core statements. Any agent with shell access connects via a pinned toolchain (elan, Lean, Mathlib) and submissions are verified server-side.
Figure 5: Decomposition graph for the Sensitivity Conjecture proved in Huang (2019) . Each dependent theorem node represents a child lemma, with an extra assisting lemma showing that A corresponds to the adjacency matrix. They are connected by a proof-sketch (purple box).Figure 6: The milestone list for the Sensitivity Conjecture mission. Each milestone pairs an authoritative statement, transcribed from the source proof, with the platform theorem the captain has attested as its canonical formalization. The two milestones shown are precisely Cauchy’s interlace theorem and the spectrum of A_{n} from Section 4.2 , here linked to the proved theorems cauchy_interlacing
Results
Between mid-June and end of July 2026, contributors closed several missions, including the exact matrix completion result, the Sipser–Gács–Lautemann theorem, and two textbooks. The largest Prove2Me mission produced 151K lines of Lean with 6 agents on consumer subscriptions, comparable in size to a centralized 30,000-agent swarm that produced 130K lines under metered API billing.
Figure 7: A snapshot of discussion under the exact matrix completion mission. Two different agents are sharing their latest progress.
Mission
LOC
Cost
Agents
Days
Algebraic Combinatorics (swarm)
130K
$100,000
30,000
7
Exact Matrix Completion
81K
$600
9
16
Sipser–Gács–Lautemann
55K
$400
3
8
Bandit Algorithms
151K
$400
6
13
Intro to Linear Optimization
17K
$200
4
7
Missions completed on Prove2Me (LOC of Lean source), with a centralized swarm for context.