arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15507cs.CLcs.LG

语言模型是否一致地编码了当前年份?

Do Language Models Consistently Encode the Current Year?

Suze van Adrichem, Aditi Bhaskar, Diyi Yang, Christopher Potts, Jing Huang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探究语言模型对当前年份的编码一致性,设计两项不同任务,发现关联与声明年份的编码机制不同,现有修改方法无法同时调整两者,表明当前年份未被一致编码。

中文摘要 AI 辅助

当前时间的一致概念对时间推理至关重要,但语言模型如何表示当前时间尚未得到充分理解。我们贡献了两项在概念上截然不同的探测当前年份的任务:关联任务,从动词时态推断当前年份;声明任务,直接查询当前年份。两项任务均将指令微调语言模型的当前年份估计控制在其训练后数据截止年份的1年范围内。对于基础模型,关联任务的预测可作为预训练数据截止年份的良好代理,13个模型的平均误差仅为10个月。然而,它们的内部机制存在差异:关联任务使用类似事实回忆的机制,而声明任务缺乏一致的因果路径。这种差异给语言模型更新当前年份带来了挑战。提示、SFT(监督微调)或权重编辑均无法成功同时调整关联年份和声明年份。提示可更新声明年份(在351个目标年份上成功率达94.6%),但几乎未改变关联年份(成功率仅1.7%)。年份偏移的SFT也未能调整关联年份,8个模型中仅有1个匹配目标年份。权重编辑虽对两项任务各自有效,但无法在两者间泛化。总体而言,我们的结果表明当前年份未在语言模型中被一致编码:关联概念深深植根于预训练学习的语言结构中,使用不同的因果机制,且难以被用于轻松调整声明概念(该概念学习自训练后阶段)的相同修改所改变。

英文摘要

A consistent concept of the current time is important for temporal reasoning, yet how language models represent the current time is not well understood. We contribute two tasks that probe the current year in conceptually distinct ways: an associative task, which infers the current year from verb tense, and a declarative task, which directly queries for the current year. Both tasks estimate current years within one year of the post-training data cutoff of instruction-tuned language models. For base models, predictions on the associative task serve as a strong proxy for the pre-training data cutoff, with an average error of only 10 months across 13 models. However, their internal mechanisms diverge: the associative task uses mechanisms similar to factual recall, while the declarative task lacks consistent causal pathways. This divergence poses a challenge for updating the current year in language models. None of prompting, SFT, or weight editing succeed in shifting the associative and declarative years simultaneously. Prompting updates the declarative year (94.6% success across 351 target years) but leaves the associative year nearly unchanged (1.7% success). Year-shifted SFT also fails to shift the associative year, matching the target year in only one of eight models. Weight editing, while effective for both tasks individually, does not generalize across both. Overall, our results show that the current year is not consistently encoded in language models: The associative notion, deeply ingrained in linguistic structures learned in pre-training, uses different causal mechanisms and resists the same modifications that easily shift the declarative notion learned in post-training.

发表机构

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑