What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models
Leonel Aguilar
cs.LG
Sep 26, 2026 · v1
TL;DR
The appendix's grouping, monotonicity and scale-invariance results for the removal-value score were formalized in Lean 4 with an AI assistant; the formalizations are in the supplementary material.
Abstract
When adapting pre-trained models through fine-tuning, freezing weights alone might not preserve performance, as updates elsewhere can change the inputs to the frozen core, ultimately affecting overall performance. We first analyse the case where a selected frozen core can be isolated and propose removal-value, a capacity-based score that approximates HOPE's removal cost averaged over removal orders. We show that in VGG-8, cutting paths from trainable neurons into a frozen core makes selection using this score useful: $70\%$ frozen preserves $5.22\pm0.51$ percentage points more old-task accuracy than DEFT at similar new-task accuracy. In transformers, shared residual streams leave paths into frozen neurons open. For this case, we derive drift-value, a forward-only proxy for the output disturbance from updating each weight entry under a local update model. In language models, at 40 epochs, this policy exceeds adapted Wanda and RIA freezing scores in settings with substantial retention loss, while its differences from Fisher remain unresolved. After 160 epochs on Qwen2.5-1.5B, it retains $0.0433\pm0.0102$ more than static Fisher. In DINOv3 vision-transformer adaptation to point clouds, drift-value retains $0.440$ image accuracy versus $0.187$ for a random mask of the same count. These results motivate Guarded Freezing: select by removal-value when incoming paths are cut, and by drift-value when they remain.
Problem
When fine-tuning pretrained models, freezing weights alone may not preserve old-task performance, because updates elsewhere can change the inputs to the frozen core. The paper asks which parameters to freeze, depending on whether paths from trainable parameters into the frozen core are cut or left open.
Approach
For isolated cores, it proposes removal-value, a capacity-based score that approximates HOPE's removal cost averaged over removal orders. For exposed cores, such as transformers with shared residual streams, it derives drift-value, a forward-only proxy for the output disturbance caused by updating each weight entry. Both scores are evaluated on VGG-8 transfer from CIFAR-100 to SVHN, on six language models learning synthetic facts, and on DINOv3 adapted to point clouds. Theoretical properties of removal-value (grouping, monotonicity, scale-invariance) are formalized in Lean 4.
Results
On VGG-8 at 70% frozen with incoming paths cut, removal-value retains 5.22±0.51 percentage points more old-task accuracy than DEFT at similar new-task accuracy. After 160 epochs on Qwen2.5-1.5B, drift-value retains 0.0433±0.0102 more than static Fisher, while its 40-epoch differences from Fisher remain unresolved. In DINOv3, drift-value retains 0.440 image accuracy versus 0.187 for a random mask of the same count.
| Policy | Image retention |
|---|
| Attention freezing only | 0.097 |
| Random mask (35.2%) | 0.187 |
| Wanda score | 0.233 |
| SSU score | 0.344 |
| Drift-value (35.2%) | 0.440 |
DINOv3 point-cloud adaptation: image retention under each freezing policy (pretrained image accuracy 0.728)