arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

零差距并非恢复:分层逐题概率评估与基准污染的逐步缓解

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Ruijie Hou, Yueyang Jiao, Zhao Wang, Yingming Li

arXiv 2608.07341首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对基准测试数据泄露导致的模型评估偏差问题,提出SA-PPG指标与RailCap方法,揭示现有缓解策略的恢复效果被高估,RailCap可实现更低的污染程度。

AI 中文摘要

公共基准的测试数据不可避免地会泄露到预训练语料库中,一旦被模型 memorize(记忆),就会夸大其评估分数。污染缓解评估会干预解码过程以抑制记忆并恢复受污染模型的真实能力,但目前流行的指标G-AP(Aggregate Performance的差距)存在缺陷:离散的正确/错误读数无法表征逐题性能,先平均再求差会导致过度抑制与抑制不足相互抵消,而统一的逐题权重会促使策略将求解概率推向干净模型的高频值。我们提出SA-PPG(Stratified Aggregate of Per-question Probability Gaps):通过采样估计每个问题的求解概率,将其与干净模型的逐题求解概率作差,并在干净模型求解概率定义的组内进行聚合。现有缓解策略首先估计污染所在位置,然后针对该估计进行操作,因此其准确性仅取决于估计的正确性。RailCap则在生成过程中判断污染:每当样本回退到贪心轨迹时,下一个轨迹 token(令牌)会被限制为亚军选项,累积抑制直到响应分布足够分散。在多个受污染模型和基准上的实验显示,SA-PPG表明先前策略的恢复效果被大幅高估,而RailCap实现了最低的SA-PPG值。

英文摘要

Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability, but its prevailing metric, the \textbf{G-AP} (\textbf{G}ap of \textbf{A}ggregate \textbf{P}erformance), is flawed. Discrete correct/incorrect readouts cannot characterize per-question performance, averaging before differencing lets over- and under-suppression cancel out, and uniform per-question weighting invites strategies to push solve probabilities onto the clean model's high-frequency values. We propose \textbf{SA-PPG} (\textbf{S}tratified \textbf{A}ggregate of \textbf{P}er-question \textbf{P}robability \textbf{G}aps): estimate each question's solve probability by sampling, difference it against the clean model per question, and aggregate within groups defined by the clean model's solve probability. Existing mitigation strategies first estimate where contamination lies and then operate on the estimate, so they are only as correct as the estimate. \textbf{RailCap} instead judges contamination during generation: whenever a sample falls back onto the greedy trajectory, the next trajectory token is capped to the runner-up, accumulating suppression until the response distribution becomes sufficiently dispersed. Across multiple contaminated models and benchmarks, SA-PPG reveals that prior strategies' restoration is substantially overestimated, while RailCap attains the lowest SA-PPG.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑