Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.