发表机构
Johannes Gutenberg University Mainz; Universidad Iberoamericana; University of Colorado Boulder(美因茨约翰内斯·古腾堡大学; 伊比利亚美洲大学; 科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现系统提示中隐藏的当前日期会导致大语言模型性能波动,影响评估可复现性,需制定谨慎的评估协议。
AI 中文摘要
可复现性对于科学研究至关重要,然而先前的研究表明,大语言模型的输出会因硬件和批处理方式的不同而变化。我们识别出一个被忽视的因素:系统提示中隐藏注入的当前日期,该日期用户无法控制且每天都会变化。在9个近期的大语言模型和6个数据集上,涵盖多项选择问答(MCQA)、数学推理、代码生成和机器翻译,模型性能仅随当前日期变化,MCQA上的差异高达6%,数学推理上为14%,代码生成上为7%,机器翻译上为2.84 BLEU。模型排名也会发生变动,影响排行榜。这种日期效应超过了其他非确定性来源,如批大小和数值精度。标准的提示技术——思维链和少样本提示——并未降低敏感性;思维链甚至放大了这种效应。我们的发现强调了需要谨慎的评估协议,以确保大语言模型研究中的可复现性和公平比较。
英文摘要
Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.
CommentsAccepted to AACL 2026 (Main)