An Accuracy–Information Tradeoff for Loss-Difference Conditional Mutual Information
Hazar Yueksel
cs.LG
Oct 6, 2026 · v1
cs.IT stat.ML
TL;DR
The main results are machine-checked in a Lean 4 formalization using Mathlib, supplied as ancillary files and mapped to declarations in an appendix.
Abstract
Loss-difference conditional mutual information (ld-CMI) uses the smallest of the standard observations in the supersample hierarchy of generalization bounds: it measures what a learner's loss differences reveal about which candidate of each pair it was trained on. Accuracy is known to force information into the model; data processing does not carry such lower bounds to losses. We show, by bounding three moments of the loss differences, that accuracy also forces ld-CMI. For linear predictors with a smooth convex loss of nonzero slope at zero, such as the logistic loss, plus a regularizer whose curvature and growth are both of power $r\ge2$, on product distributions over a scaled sign cube in dimension at least linear in $n$, every proper learner with expected excess risk at most $\varepsilon$ on these distributions at the optimal sample size $n\asymp\varepsilon^{-2+2/r}$ has worst-case ld-CMI of order $n$ bits, and $Θ(n/(1+(τ/\varepsilon)^2))$ bits under Gaussian noise of standard deviation $τ$ on the loss differences. The same holds without a regularizer, at $n\asymp\varepsilon^{-2}$. Consequently, range-scaled ld-CMI bounds cannot vanish on these distributions, although every proper learner's generalization gap is $O(n^{-1/2})$. We also show that model-level information does not determine noisy loss-difference information, and that the growth, slope and dimension conditions are needed, the last up to a logarithm.
Problem
It is known that accuracy forces information into a learned model. Data processing does not carry such lower bounds over to loss-difference conditional mutual information (ld-CMI), the smallest observation in the supersample hierarchy of generalization bounds. The question is whether accuracy also forces ld-CMI, and how much observation noise removes it.
Approach
The setting is linear predictors with a smooth convex loss of nonzero slope at zero and a regularizer of power r≥2, on product distributions over a scaled sign cube. Three moments of the loss differences are bounded for every accurate learner, and a transfer theorem turns these moment bounds into two-sided information bounds at every Gaussian noise level. Counterexamples show which hypotheses are needed. The main results are formalized in Lean 4 with Mathlib.
Results
At the optimal sample size, every ε-accurate proper learner has worst-case ld-CMI Θ(n/(1+(τ/ε)²)) bits. Range-scaled ld-CMI bounds therefore cannot vanish on these distributions, even though every proper learner's generalization gap is O(n^{-1/2}). Model-level information does not determine noisy loss-difference information, and the growth, slope and dimension conditions are needed, the last up to a logarithm.