Title: From Overparameterization to Sparsity-Aware Optimization
Abstract: Overparameterization lies at the heart of the dominant paradigm in AI development: collecting large datasets to train even larger neural networks. One consequence is the substantial resource demand, which makes cutting-edge development accessible to only a small number of labs. In this talk, we ask whether it really has to be like this. While sparse solutions are known to exist in theory, overparameterization appears to play a crucial role in optimization success, a phenomenon that has been linked to the implicit bias of the induced training dynamics. Motivated by the goal to transfer the benefits of overparameterization to sparse training, we introduce a framework for controlling implicit bias and exploiting it to induce sparsity and to address a central challenge: learning parameter sign flips. Building on our insights, we discuss how adjusting the relative learning speeds of layers can improve sign learning and motivate the design of sparsity-aware optimizers.