Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
Sebastian Bobadilla-Suarez, Bob Suh, Ryan Fortin
cs.LG
Sep 28, 2026 · v1
cs.AI cs.SE
TL;DR
Algebraic steps of the saturation, orchestration-ceiling and majority-vote results are machine-checked in Lean 4 with Mathlib, with no sorry; Lean refuted one conjectured strict inequality.
Abstract
An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape only if that set expands. Rewriting scaffolding (tools, verifiers, decomposition) expands what an agent reaches without touching a weight, so frozen weights buy an eventual ceiling but no stationarity along the way. The criterion also separates three regimes usually merged: search within a fixed class, test-time training that raises the ceiling itself, and scaffold rewriting between them. Audit the scaffold, not the checkpoint. The same ceiling binds sideways. Best-of-$k$ orchestration realizes the best worker's ceiling exactly: width buys rate, not budget. Re-consulting a fixed pool has a horizon computable in advance, decided by the pool alone, and the one arrangement that would beat it, a weighted vote, needs diversity real workers lack: on 30 same-family workers the failure overlap sits at its maximum, and a majority fails 23/55 (42%) of tasks. We obtain the criterion by reading refinement as gradient boosting on the residual error between draft and target, a patch or git diff, and then measuring where that reading breaks: patches compose instead of standing beside each other to be voted on, and failures overlap. What we measure is saturation. Per-round improvement decays toward zero on SWE-bench, and churn decays geometrically across 401 production sessions, a shape shared with a pre-AI human baseline that establishes the regime without identifying its cause. Both breaks are engineering choices rather than laws about code, so together they specify a harness worth building.
Problem
The question is when iterative self-modification by agentic coding systems runs out of room, and whether frozen model weights are sufficient evidence that it has. The study also asks how orchestration and refinement loops should be designed.
Approach
Refinement is modeled as gradient boosting on the residual between a draft and the target code, represented as a patch. This yields a refinement game and a stationarity dichotomy: returns strictly diminish when the reachable edit set is fixed. Orchestration ceilings and majority-vote bounds are derived, and the algebraic steps are machine-checked in Lean 4 with Mathlib, including product-measure independence. Empirical tests use a mini-swe-agent on 55 SWE-bench Lite tasks and 401 production sessions.
Results
Best-of-k orchestration exactly achieves the best worker's ceiling, and Lean found a counterexample to a conjectured strict inequality. Worker failure overlap is maximal: a majority of 30 workers fails 23/55 tasks. Per-round improvement decays toward zero, and churn decays geometrically, matching a pre-AI human baseline.