发表机构
Department of Computer Science Cornell University(计算机科学系 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出连续查询有限记忆语言模型CO-LMLM,通过将知识库的键与文本知识值配对,以低成本生成灵活向量查询,整合检索知识。经多数据集预训练和多模型规模测试,在困惑度和事实精度上优于同类模型。
AI 中文摘要
有限记忆语言模型(LMLMs)在预训练期间将事实知识外化到知识库(KB)中,而非存储在权重里。生成时,模型按需从KB获取知识。我们提出连续查询LMLM(CO-LMLM),其KB将连续键与文本知识值配对。CO-LMLM以低成本生成灵活向量查询,还将人类可读且可归因的检索知识整合到生成中。通过在Wikipedia和FineWeb-Edu上预训练并在多个模型规模下测试,CO-LMLM在困惑度和事实精度上均优于先前的LMLMs和普通LLMs。
英文摘要
Externalizing knowledge in LLM pre-training is a promising avenue to achieve higher performance at smaller scales, control knowledge use, and overall increase model transparency. We propose continuous-query limited memory language models (Co-LMLM), an LLM that interleaves flexible vector retrieval queries with next-token predictions, and is pre-trained to copy knowledge returned from the KB, rather than memorize it. Co-LMLM is pre-trained with a scalable approach that jointly trains a knowledge-externalizing LLM, induces its knowledge base, and learns an expressive continuous retrieval mechanism. Across pre-training at multiple model scales, Co-LMLM outperforms prior knowledge-externalizing and vanilla LLMs in both perplexity and factual precision. At 360M scale, this includes lower perplexity than models pre-trained on 40$\times$ more data, and SimpleQA-verified performance that is in line with gpt-4o-mini and higher than Claude Sonnet 4.5.
Commentspreprint. Project page: https://lil-lab.github.io/co-lmlm-web/