Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
AI Digest - ArXiv AI
Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning
Source: ArXiv AI