[论文] Knowledge Acquisition During Pre-training? Large Language Models Learn…

## 论文概要 **研究领域**: NLP **作者**: Joseph Lee, Yidi Huang, D...

论文概要

研究领域: NLP 作者: Joseph Lee, Yidi Huang, Dokyoon Kim 发布时间: 2026-09-06 arXiv: 2509.04288

中文摘要

关于大语言模型(LLM)在预训练期间如何获取知识,仍存在理解空白。我们假设辅助视角——即知识的重新表述——对学习具有因果帮助。我们设计了对照实验来分离这一效应。首先,我们确认了重复对知识获取是必要的,并明确改述仅在较小批次大小时有帮助。其次,在固定token预算下,将token从文档重复分配到辅助视角能提升学习效果,反直觉的是,这对事实回忆也成立。第三,辅助视角的有效性不依赖于生成它们的教师模型的强度。第四,我们识别了在有先验知识缺口时辅助学习的知识形式:情境性和基础性知识。最后,我们通过逐层偏差和压缩机制地检验了这些效应的表现方式。综合来看,我们的发现表明,辅助知识表示(在大规模预训练语料中自然产生)是预训练成功的关键因素,并为数据多样性为何重要提供了合理解释。

原文摘要

Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Final…

自动采集于 2026-09-07

#论文 #arXiv #NLP #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1