可验证腐败预算:自适应操纵下的任意时刻有效排行榜声明
Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
浏览论文内容
中文总结 AI 辅助
针对AI模型排行榜易受自适应操纵的问题,提出了带数学公式的可验证腐败预算方法,该方法能在任意时刻保证排行榜声明的正确性,在Chatbot Arena投票实验中表现出良好的抗操纵能力。
中文摘要 AI 辅助
AI模型的公开排行榜会被持续读取,攻击者可以看到每一条已发布的排名。投票操纵、私有变体的选择性披露以及基准污染都可能改变排名。现有保证假设记录是真实的,或对每一步的腐败进行了限制,但进行批量腐败的攻击者可以规避这些限制。我们引入了可验证腐败预算,这是一种在$t$条记录后计算并随每一条成对声明一同发布的容差$\widehat{B}_t$。以至少$1-\alpha$的概率,在所有时刻,该声明要么是正确的,要么有超过$\widehat{B}_t$条记录被腐败。它适用于每一条证书都查看、且预算无限制的攻击者。伪造记录和被查看后被更改的记录需要不同的证书:针对翻转已查看投票的攻击者,伪造证书失效的概率趋近于1;而每条记录收费约两倍的证书仍然有效,即使是面对能预见未来、且赌注恒定的攻击者,在任何级别上都不存在更小的有效收费。可验证预算的增长速度几乎与任何有效方法允许的速度一样快:在胜率$\frac{1}{2}+\delta$的情况下,每条新记录会为该声明可承受的伪造记录数量增加近$2\delta$(对应$\delta$次翻转)。发布$V$个私有变体中的最佳者,仅会产生$\log V$量级的增长成本。在对180万条Chatbot Arena投票的重放实验中,数百条操纵投票就会使标准置信区间验证出错误的排序,而我们的方法仍保持有效。在真实投票中,我们的证书显示,明显区分的模型可承受约2000条伪造投票。
英文摘要
Public leaderboards for AI models are read continuously, and attackers can see every published standing. Vote rigging, selective disclosure of private variants, and benchmark contamination can each move a ranking. Existing guarantees assume genuine records or bound the corruption per step, which an attacker who corrupts in bursts evades. We introduce the certified corruption budget, a tolerance $\widehat{B}_t$ computed after $t$ records and published with each pairwise claim. With probability at least $1-α$, simultaneously at all times, the claim is correct or more than $\widehat{B}_t$ records were corrupted. It holds against attackers who watch every certificate, with no bound on their budget. Forged records and records altered once seen require different certificates: the certificate for forgeries fails, with probability approaching one, against an attacker who flips votes it has seen, while one that charges roughly twice as much per record remains valid, with constant bets even against attackers who see the future, and no smaller charge is valid at every level. The certified budget grows nearly as fast as any valid method allows: with a win fraction $\frac{1}{2}+δ$, each new record adds close to $2δ$ to the number of forged records the claim can withstand ($δ$ flipped). Publishing the best of $V$ private variants costs only an amount growing like $\log V$. In replays on 1.8 million Chatbot Arena votes, a few hundred rigged votes make standard confidence intervals certify false orderings, while ours stays valid. On real votes, our certificate shows that clearly separated models withstand about 2,000 forged votes.