arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需预言机的排序:来自离线多智能体日志的偏差感知交互排序选择

Rank Without an Oracle: Deviation-Aware Interaction-Rank Selection from Offline Multi-Agent Logs

Xiangwu Wang, Chengwei Cao, Hongyuan Tang

arXiv 2609.08358首次发表:更新:

发表机构

University of Hong Kong; University of California, San Diego; Carnegie Mellon University(香港大学; 加州大学圣地亚哥分校; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对离线多智能体日志,提出偏差感知的交互排序选择方法SIRV,通过公共分布并集校准和弃权机制,提供可认证的模型选择,并显著降低CCE差距证书。

AI 中文摘要

离线多智能体收益模型是在日志分布下估计的,但被用于由学习到的解和单边偏离所诱导的分布上。因此,标准的留出损失可能偏向于能很好预测日志行为但扭曲战略激励的交互类别。我们针对具有已知日志分布的有限博弈引入了选择性交互排序验证(SIRV)。训练集拟合嵌套收益模型,并构建所有候选部署分布和单边替换分布的公共并集;独立的校准集在此同一并集上评估每个候选。SIRV返回最小的排序,其同时上最坏目标风险在最佳上分数的容差范围内,并在声明目标不受支持或估计过于不精确时弃权(不执行)。一个公共覆盖事件产生有限候选目标风险界和候选特定的粗相关均衡(CCE)差距证书。我们还分离出一个精确的两点非支持不可识别性结果。在一项受控因子研究中,每个族有2,048个独立博弈,经验-伯恩斯坦界相对于霍夫丁界在公共回报上将中位CCE差距证书减少了42.5%,在支持回报上减少了1.36点。在配对排序错误设定和单独生成的拥塞族中,SIRV-EB回退规则相对于ID-Mean降低了平均真实候选选择CCE遗憾,同时保留了博弈级损失。在$N=3,5,8$的384个博弈中,ID-Mean相对平均CCE遗憾效应保持为正,而在弱覆盖下认证回报急剧下降。这些结果将可认证的模型选择与普遍的战略改进区分开来。

英文摘要

Offline multi-agent payoff models are estimated under a logging distribution but used on distributions induced by learned solutions and unilateral deviations. Standard held-out loss can therefore favor an interaction class that predicts logged play well while distorting strategic incentives. We introduce Selective Interaction-Rank Validation (SIRV) for finite games with known logging distributions. A training split fits nested payoff models and constructs a common union of all candidate deployment and unilateral-replacement distributions; an independent calibration split evaluates every candidate on this same union. SIRV returns the smallest rank whose simultaneous upper worst-target risk is within tolerance of the best upper score, and abstains when a declared target is unsupported or too imprecisely estimated. A common coverage event yields a finite-candidate target-risk bound and a candidate-specific coarse correlated equilibrium (CCE) gap certificate. We also isolate an exact two-point off-support non-identifiability result. In a controlled factorial study with 2,048 independent games per family, empirical-Bernstein bounds reduce the median CCE-gap certificate by 42.5% relative to Hoeffding bounds on common returns, with a 1.36-point reduction in supported return. Under paired rank misspecification and in a separately generated congestion family, the SIRV-EB fallback rule lowers mean true candidate-selection CCE regret relative to ID-Mean, while retaining game-level losses. Across 384 games at $N=3,5,8$, ID-Mean-relative mean CCE-regret effects stay positive while certified return falls sharply under weak coverage. These results separate certifiable model selection from universal strategic improvement.

Comments18 pages, 9 figures, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑