arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于大语言模型时序评估与知识更新的合成世界

Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs

Jonathan Zheng, Zirui Shao, Alan Ritter, Wei Xu

arXiv 2609.00184首次发表:更新:

发表机构

Georgia Institute of Technology; Zhejiang University(佐治亚理工学院; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型知识过时问题,提出ParallelEvents基准与Synapse训练框架,实现无人工标注的可扩展知识插入,性能较现有方法提升14.23%。

AI 中文摘要

大语言模型(LLMs)依赖静态预训练语料,导致其知识随时间变得过时。现有知识编辑评估方法要么存在快速污染问题,要么依赖与现有刚性知识冲突的反事实编辑。本研究提出一种基于合成、模拟驱动的框架,用于研究LLMs中的知识插入。我们引入ParallelEvents基准,该基准包含虚构但真实的未来世界,可生成连贯事件轨迹以实现可控评估,避免污染同时保持一致性。基于此数据集,我们开发Synapse训练框架,利用模型生成的数据通过中间训练和指令微调更新模型参数。该合成流程可实现可扩展的知识整合,无需昂贵的人工整理数据。实验表明,Synapse的性能比现有方法高出14.23%,证明基于模拟的合成训练可实现稳健且连贯的知识插入。

英文摘要

Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactual edits that conflict with rigid existing knowledge. In this work, we propose a synthetic, simulation-driven framework for studying knowledge insertion in LLMs. We introduce {\sc ParallelEvents}, a benchmark of fictional yet realistic future worlds that generates coherent event trajectories for controlled evaluation, avoiding contamination while preserving consistency. Building on this dataset, we develop {\sc Synapse}, a training framework that uses model-generated data to update model parameters via mid-training and instruction tuning. This synthetic pipeline enables scalable knowledge integration without costly human-curated data. Empirically, {\sc Synapse} outperforms existing methods by 14.23\%, demonstrating that simulation-based synthetic training leads to robust and coherent knowledge insertions.

Commentspreprint, 12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑