发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究测量LLM事实核查中的参数泄漏与事后证据泄漏,发现两者可被总体准确率掩盖,表示瓶颈能有效减少泄漏,且事后证据在AVeriTeC上虚增6.3点准确率,揭示基准测试可能高估性能。
AI 中文摘要
随着自动化事实核查在社交媒体上的规模化应用,大语言模型(LLM)的判定分数可能看起来比实际应有的更强。原因之一是评估中混入了在声明提出时无法获知的信息。有两个通道容易被混淆:预训练中记忆的结果和声明发布后检索到的证据。然而,标准基准测试很少将这两者区分开来。在本研究中,我们通过在 AVeriTeC 和 QuanTemp++ 上重建时间点证据条件并探测模型表示中编码的结果信息,对这两个通道进行了测量。我们发现了大量参数泄漏的证据,这些泄漏可能被总体准确率所掩盖,而一个简单的表示瓶颈比基于互信息的训练惩罚更有效地减少了这种未来泄漏。我们还发现,允许声明后证据使 AVeriTeC 的零样本准确率虚增了 6.3 个百分点,而在 QuanTemp++ 中该效应可忽略不计,因为检索提供的声明后证据很少。这些结果表明,当错误信息基准测试未考虑声明时实际可用的信息时,可能会高估事实核查性能。
英文摘要
As automated fact-checking scales on social media, large language model (LLM) verdict scores can look stronger than warranted. One reason is that evaluations mix in information that was not knowable at claim time. Two channels are easy to conflate: outcomes memorized in pre-training and retrieved evidence published after the claim. Yet standard benchmarks rarely separate the two. In this study we measure both channels on AVeriTeC and QuanTemp++ by reconstructing point-in-time evidence conditions and probing for outcome information encoded in model representations. We find substantial evidence of parametric leakage, that can be hidden by the aggregate accuracy, while a simple representation bottleneck reduces this future leakage more efficiently than a mutual-information-based training penalty. We also find that allowing post-claim evidence inflates zero-shot accuracy by 6.3 points in AVeriTeC while the effect is negligible in QuanTemp++, where retrieval provides little post-claim evidence. These results show that misinformation benchmarks can overstate fact-checking performance when they do not account for what information was actually available at claim time.
Comments11 pages, 5 figures