arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型回测中的时间泄漏:测量、验证与调整后得分

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

Zeyu Zhang, Bradly C. Stadie

arXiv 2608.02985首次发表:更新:

发表机构

Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出LLM回测的标准污染检测方法无效,提出利用已知截止时间和匹配干净对照组测量时间泄漏的方法,可检测并调整回测得分,澄清部分模型优势源于近期性而非真实技能。

AI 中文摘要

大语言模型(LLM)回测中污染的标准检测方法十分简单:比较训练截止时间前后的得分。我们证明这种检测方法毫无意义。四款旗舰模型在未记忆的问题上均未通过该检测:所有得分问题均在其截止时间后得到解决。原因是结构性的:模型合法地了解更多关于其截止时间附近的时间信息,因此近期性会模拟泄漏,且我们证明任何被动回测都无法将这两者与真正的技能区分开。测量(而非仅检测)需要回测之外的信息。我们以两种形式提供此类信息:已知截止时间可识别边界处的泄漏;匹配的干净对照组可在全局识别泄漏并产生泄漏调整后得分。我们还推导了泄漏隐藏的位置:它集中在让大众感到意外且在训练中被充分覆盖的结果上,且部分记忆会得到不成比例的奖励。我们通过在孪生模型中植入泄漏来针对真实值验证估计量,这些估计量可恢复注入的剂量,并在干净问题上返回空值。将其应用于前沿模型时,它们检测到一个截止时间局部的特征,且在审计的功效下限下,澄清了五款模型,其明显优势仅来自近期性。回测无需被丢弃;它们只需要一个可辩护的参考。

英文摘要

The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff. We show this check is uninformative. Four flagship models fail it on questions they cannot have memorized: every scored question resolved after their cutoffs. The reason is structural. Models legitimately know more about times near their cutoff, so recency mimics leakage, and we prove no passive backtest can separate the two from genuine skill. Measurement, not just detection, requires information from outside the backtest. We supply it in two forms. A known cutoff identifies leakage at the boundary; a matched clean control identifies it globally and yields a leakage-adjusted score. We also derive where leakage hides: it concentrates on outcomes that surprised the crowd and were well covered in training, and partial memorization is disproportionately rewarded. We validate the estimators against ground truth by planting leakage in twin models, where they recover the injected dose and return null on clean questions. Deployed on frontier models, they detect one cutoff-localized signature and, at the audit's power floor, clear five models whose apparent advantages were recency alone. Backtests need not be discarded; they need one defensible reference.

Comments12 pages main content, 45 pages in total

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑