Structural priors for data-efficient language learning
Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight initialization for multilingual language modeling. We evaluate transfer via next-token-prediction loss, weight shifts in the…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.