← All papers
First page of Double descent is the principle of least action

Double descent is the principle of least action

Congzhou M Sha

cs.LG Sep 16, 2026 · v2 cs.AI math.ST physics.comp-ph physics.data-an
The paper's theorems, including the result that the stationary-path energy is non-increasing in d and the quadratic completion identity, are verified in Lean 4 with Mathlib. The proofs are supplied as ancillary files.
The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon. We explain the phenomenon with statistical mechanics. The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature $T$, and a run that has equilibrated visits every parameter vector of a given training loss equally often, the fundamental postulate of statistical mechanics, with probability given by the Boltzmann distribution. Because training starts at an initial point and has only finite time to diffuse, it carries an effective weight decay, which makes every parameter a quadratic degree of freedom. The equipartition theorem then distributes the energy among the $d$ degrees of freedom in shares of $T/2$, so at a fixed training loss adding parameters lowers the temperature and drives the Boltzmann distribution toward the stationary path. Finally, adding parameters can only lower the $L^2$ norm of the stationary path, so a solution sampled at fixed loss is less likely to be large with increasing $d$, effectively increasing weight regularization.

Test error plotted against model parameter count falls, peaks at the interpolation threshold, and falls again (double descent). The paper seeks an elementary, pedagogical statistical-mechanics explanation of this phenomenon.

SGD training is modeled as a particle diffusing over the loss landscape at an induced temperature, with equilibrium given by the Boltzmann distribution. Finite training time from an initial point acts as an effective weight decay, so every parameter becomes a quadratic degree of freedom. Equipartition then shows that adding parameters lowers the temperature at fixed loss, pushing samples toward the stationary path. A theorem shows that adding parameters can only lower the stationary path's L2 norm. Theorems and key identities are machine-checked in Lean 4 with Mathlib, with source files provided as arXiv ancillary files.

Increasing the parameter count d beyond the interpolation threshold acts as implicit weight regularization, making large-norm solutions less likely at fixed training loss. This is illustrated on Legendre polynomial regression of noisy cubic data.