AI Agents Co-Synthesize Code and Proofs for Verified Distributed Systems
The paper introduces Inductive Deductive Synthesis (IDS), an agentic LLM system that jointly and incrementally builds an implementation alongside its formal correctness proof, learning from failed attempts to guide future strategies. While SOTA coding agents (Codex/GPT-5.4, Claude Code/Opus 4.6) solve only 2/7 distributed key-value-store verification specs, IDS reportedly achieves 7/7 in about 6.8 hours and $106 per spec—roughly 200x faster than expert human effort and 17% cheaper than existing agents—while also optimizing implementations to run up to 3x faster than published verified systems. The authors frame this as a step toward giving AI agents formal correctness guarantees that testing alone cannot provide. On Twitter, the authors highlighted the NeurIPS oral acceptance and framed the work as a leap for combining agentic coding with formal verification, positioning it alongside related efforts on specifying human intent correctly (a follow-up challenge) and other lab outputs like SPECS and BenchEvolver. Discussion so far is mostly self-promotional from the authors, with no substantive external critique yet visible in the available tweets.
Discussion: 2 tweets from 2 authors · @shulynnliu, @mertcemri