← All papers
First page of LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization

LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization

Lazar Milikic, Simon Guilloud, Khanh Nguyen, Viktor Kuncak

cs.AI Jun 26, 2026 · v1 cs.LG cs.LO
Presents an LLM agent runtime that turns mathematical papers into buildable Lean 4/Mathlib projects, with a queue-based prover and the LeanProbe cached checker.
We present and evaluate LeanFlow, an LLM agent system specialized for translating mathematical papers into buildable Lean projects. Recent verifier-in-the-loop systems show that large formal artifacts can be produced, but it remains unclear which runtime mechanisms affect completion, auditability, or efficiency in document-to-project formalization. We study this question through case studies on two previously unformalized mathematical papers in number theory and measure theory, using model, proof-workflow, and toolset ablations with Kimi2.6 and GPT5.5; we report task outcome, API calls, input tokens, and output tokens. With Kimi2.6, the full workflow completes both document-level projects within the 2000-call budget, while no-queue variants reach the budget limit; with GPT5.5, all document-level variants complete, and the full workflow has the lowest or tied-lowest input-token cost on both sources. As complementary calibration, LeanFlow reaches 75.7% BEq+ on the PFR slice of RLM25 and solves all five ICML 2026 AI for Math TCS challenge projects in our GPT5.5 runs.

Verifier-in-the-loop systems can produce large Lean artifacts from mathematical documents. It is unclear which runtime mechanisms affect completion, auditability, and efficiency when a whole paper is formalized into a buildable project.

LeanFlow runs a deterministic source preflight and builds a blueprint mapping source spans to planned Lean declarations. A statement/source gate checks faithfulness before proof search. A programmatic workflow manager then assigns one proof obligation at a time, keeps failed-attempt memory, and accepts an edit only after Lean/Lake verification and a hygiene scan for sorry and axioms. LeanProbe supplies low-latency cached checks, and ablations remove the queue and swap the full toolset for CLI-only access, using Kimi-K2.6 and GPT-5.5.

With Kimi-K2.6, the full workflow completed both previously unformalized papers (a number-theory paper on Pythagorean polynomials and a measure-theory paper on Cramer–Wold) within the 2000-call budget, while no-queue variants exhausted the budget. With GPT-5.5, all document-level variants completed, and the full workflow had the lowest or tied-lowest input-token cost. LeanFlow reached 75.7% BEq+ on the PFR slice of RLM25 and solved all five ICML 2026 AI for Math TCS challenge projects.

SourceWorkflowOutcomeCallsIn tok.
PythagoreanFullsuccess104346.9M
PythagoreanNoQ+Toolsfailure2000160.2M
Cramer–WoldFullsuccess127866.0M
Cramer–WoldNoQ+Toolsfailure2000127.8M
Kimi-K2.6 document-level ablations (excerpt)
WorkflowProof succ.BEqBEq+Calls
LeanFlow81.2%70.8%75.7%3541
terminal-only agent80.6%68.1%72.9%4053
RLM25-PFR autoformalization with GPT-5.5