A Spectral Theory of Grokking: Weight Decay induces Feature Learning
In grokking an early fit to the training data separates from a much later improvement in generalization. During this delay, training can move from a fixed neural tangent kernel (NTK) regime to one in which task-relevant kernel eigendirections continue to evolve. We provide a quantitative theory for how this transition from lazy to rich learning can produce delayed generalization. For homogeneous…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.