← All papers
First page of Removing spurious minima for planar features by skip connections

Removing spurious minima for planar features by skip connections

Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho

cs.LG Oct 1, 2026 · v1 cs.AI math.OC
Main theorems on ReLU teacher-student loss landscapes are formalized in Lean, with each proposition and theorem linked to a Lean artifact (Lean_6–Lean_12).
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher–student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.

The structure of loss landscapes for shallow bias-free ReLU networks with Gaussian inputs in the teacher-student setting is only partially understood. The open question is when overparameterization or a learned linear skip connection removes spurious local minima.

The Gaussian population loss is rewritten in mass-direction coordinates using a residual representation. Feature-learning and zero-mass arguments reduce planar teacher networks to an angular analysis on the circle, using interlacing of student and teacher lines. A computer-assisted strong-convexity certificate is used to construct spurious minima for plain ReLU networks. Key results are accompanied by Lean formalizations.

With a learned linear skip and planar positive teachers, every local minimum with non-negative student masses is an exact fit once the student width is at least the teacher width. Without the skip, a width-3 teacher in 2D has spurious minima at every student width of at least 3. In 2D, critical points have at most 2m distinct student directions, and the benignity result transfers to r-stable empirical minima.