Layer rotation: a surprisingly simple indicator of generalization in deep networks?

Carbonnelle, Simon;De Vleeschouwer, Christophe
(2019) 36th International Conference on Machine Learning, Identifying and Understanding Deep Learning Phenomena workshop — Location: Long Beach (9.June.2019)

Files

ICML2019-Simon-compressed.pdf
  • Open Access
  • Adobe PDF
  • 375.66 KB

Details

Authors
Abstract
Our work presents empirical evidence that layer rotation, i.e. the evolution across training of thecosine distance between each layer’s weight vector and its initialization, constitutes an impressively consistent indicator of generalization performance. Compared to previously studied indicators of generalization, we show that layer rotation has the additional benefit of being easily monitored and controlled, as well as having a network independent optimum: the training procedures during which all layers’ weights reach a cosine distance of 1 from their initialization consistently outperform other configurations -by up to 20% test accuracy. Finally, our results also suggest that the study of layer rotation can provide a unified framework to explain the impact of weight decay and adaptive gradient methods on generalization.
Affiliations

Citations

Carbonnelle, S., & De Vleeschouwer, C. (2019). Layer rotation: a surprisingly simple indicator of generalization in deep networks? Proceedings of the 36th International Conference on Machine Learning, 1(1), 8. https://hdl.handle.net/2078.5/126782 (Original work published 2019)