发表机构
IIIS, Tsinghua University(清华大学交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AV-AIVAT将AIVAT与置信序列结合,在非完全信息博弈中实现任意时间有效停止,成本降低74倍,可高效评估智能体并向第三方提供可审计的验证依据。
AI 中文摘要
判断两个智能体中哪个更强需要进行博弈直到技巧胜过运气,而每一局博弈都会耗费资金、模型推理时间或专家时间。由于所需博弈局数未知,固定预算评估要么在结果确定后仍持续投入资源,要么在能区分智能体前就停止;而使用普通置信区间的朴素可选停止会使指定的置信水平失效。我们实现了这类评估在证据充足时立即停止,同时保持有效性。行动知情价值评估工具(AIVAT)通过条件均值校正降低非完全信息博弈的方差,在涵盖71439局一对一无限注德州扑克(HUNL)手牌的15种大语言模型(LLM)智能体配置中,其中位数降低幅度达54倍,但AIVAT未规定何时停止。我们将AIVAT与持续监测的置信序列(CSs)结合,形成任意时间有效AIVAT(AV-AIVAT),其在线价值模型仅从过往博弈中学习,避免对局自身结果进行校正。在名义95%水平、目标精度为±1大盲注的条件下,原始结果在渐近置信序列(AsympCS)下停止所需的手牌数中位数是经AIVAT校正结果的74倍。精确有限样本验证使用经验伯恩斯坦置信序列(EB-CS),该序列需要对校正后收益提供独立验证的界;我们针对德州扑克里德克(Leduc hold'em)博弈从结构上建立了该界,并确定了由CS的赌注上限和该界设定的宽度下限,该下限决定了多少方差增益可转化为更早停止;描述性HUNL EB-CS实验显示,停止时间比的中位数为1.37倍。AV-AIVAT将方差降低转化为高效、可审计的早期停止,同时分离渐近筛选与精确验证,使评估能在证据充足时立即停止,并向第三方提供在该停止时间重新验证裁决所需的全部内容。
英文摘要
Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while naive optional stopping with an ordinary confidence interval invalidates the stated level. We make such an evaluation stop as soon as its evidence suffices, with the guarantee intact. The Action-Informed Value Assessment Tool (AIVAT) reduces variance in imperfect-information games through conditional mean-zero corrections, by a median $54\times$ across 15 LLM agent configurations spanning 71,439 paired Heads-Up No-Limit Hold'em (HUNL) hands, but does not say when to stop. We combine AIVAT with continuously monitored Confidence Sequences (CSs) into anytime-valid AIVAT (AV-AIVAT), whose online value model learns only from past games so that no game scores its own correction. At the nominal 95\% level and a target precision of $\pm1$ Big Blind, raw outcomes need a median $74\times$ as many hands as AIVAT-corrected outcomes to stop under the Asymptotic CS (AsympCS). Exact finite-sample certification uses the Empirical-Bernstein CS (EB-CS), which needs an independently justified bound on corrected payoffs. We establish such a bound structurally for Leduc hold'em and characterize a width floor set by the CS's bet cap and that bound, which governs how much of a variance gain becomes earlier stopping; the descriptive HUNL EB-CS runs show a median $1.37\times$ stopping-time ratio. AV-AIVAT turns variance reduction into efficient, auditable early stopping while separating asymptotic screening from exact certification, so an evaluation can stop the moment its evidence suffices and hand a third party everything needed to recheck the verdict at that very stopping time.
Comments34 pages, 5 figures