arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为模型标注日期:系统提示中的隐藏日期影响大语言模型评估

Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

Mario Sanz-Guerrero, Minh Duc Bui, Manuel Mager, Katharina von der Wense

arXiv 2609.36931首次发表:更新:

发表机构

Johannes Gutenberg University Mainz; Universidad Iberoamericana; University of Colorado Boulder(美因茨约翰内斯·古腾堡大学; 伊比利亚美洲大学; 科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现系统提示中隐藏的当前日期会导致大语言模型性能波动,影响评估可复现性,需制定谨慎的评估协议。

AI 中文摘要

可复现性对于科学研究至关重要,然而先前的研究表明,大语言模型的输出会因硬件和批处理方式的不同而变化。我们识别出一个被忽视的因素:系统提示中隐藏注入的当前日期,该日期用户无法控制且每天都会变化。在9个近期的大语言模型和6个数据集上,涵盖多项选择问答(MCQA)、数学推理、代码生成和机器翻译,模型性能仅随当前日期变化,MCQA上的差异高达6%,数学推理上为14%,代码生成上为7%,机器翻译上为2.84 BLEU。模型排名也会发生变动,影响排行榜。这种日期效应超过了其他非确定性来源,如批大小和数值精度。标准的提示技术——思维链和少样本提示——并未降低敏感性;思维链甚至放大了这种效应。我们的发现强调了需要谨慎的评估协议,以确保大语言模型研究中的可复现性和公平比较。

英文摘要

Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.

CommentsAccepted to AACL 2026 (Main)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑