Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
Shubham Agarwal, Alexander Krentsel, Shu Liu, Mert Cemri, Audrey Cheng, Rui Meng, Tomas Pfister, Chun-Liang Li, Sylvia Ratnasamy, Aditya Parameswaran, Matei Zaharia, Ion Stoica, Mohsen Lesani
cs.AI
May 22, 2026 · v1
cs.DC cs.LO cs.PL
TL;DR
Rocq is the primary backend; the proof agent is also evaluated on Lean 4 benchmarks (VERINA, miniCodeProps), checked via lake build.
Abstract
AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage that testing alone cannot provide. Distributed systems are a prime example: properties such as consistency between reads and writes must hold under every possible interleaving of events. Mechanized formal verification can guarantee such correctness, but typically demands months to years of expert effort. As evidence, even SOTA coding agents (Codex with GPT-5.4 and Claude Code with Opus 4.6) succeed on only 2/7 distributed key-value-store specifications. In this paper, we present the first effective approach to addressing this gap, Inductive Deductive Synthesis (IDS), which jointly and incrementally synthesizes implementation and proof, and learns from failed attempts to systematically try promising strategies. Built as an agentic LLM system, IDS achieves 7/7 in about 6.8 hours and $106 per spec on average, roughly 200x faster than expert effort and 17% cheaper than SOTA agents. IDS further incorporates performance feedback into the same loop, yielding implementations up to 3x faster than published verified systems.
Problem
Coding agents cannot reliably produce formally verified distributed systems. Strong agents such as Codex (GPT-5.4) and Claude Code (Opus 4.6) each succeed on only 2 of 7 key-value-store consistency specifications. Manual mechanized verification of such systems typically takes months to years of expert effort.
Approach
Inductive Deductive Synthesis (IDS) is an agentic LLM system that synthesizes implementation and proof jointly and incrementally. Deductive Synthesis Agents decompose components using Admitted placeholders, and the Rocq type-checker checks each partial proof as it goes. An Inductive Synthesis Agent learns from failed attempts: it proposes helper lemmas when progress stalls and respawns agents with new designs at dead ends. Completed implementations are extracted to OCaml and benchmarked for performance feedback. The system is also evaluated on cross-language benchmarks in Dafny, Verus, Lean 4 and Rocq.
Results
IDS solves all 7 specifications at about 6.8 hours and $106 per spec on average. Its implementations are up to 3x faster than published verified systems. On proof benchmarks it reaches state-of-the-art results, including 176/189 on Lean's VERINA and 100/100 on miniCodeProps.
| Method | DafnyBench | miniCodeProps | Verus-Bench | CoqStoq | VERINA |
|---|
| Prior SOTA | 52/100 | 49/100 | 137/150 | 28/100 | 38/189 |
| Codex | 78/100 | 80/100 | 148/150 | 46/100 | 136/189 |
| Claude Code | 76/100 | 86/100 | 148/150 | 51/100 | 149/189 |
| IDS | 88/100 | 100/100 | 149/150 | 97/100 | 176/189 |
Pass counts on verification benchmarks (selected)