arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RSD-Poker:不完全信息博弈中残差策略的结构自适应与移位鲁棒风险-效用认证

RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games

Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Peng Zhang, Daren Zha, Jun Xiao

arXiv 2609.33669首次发表:更新:

发表机构

School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院大学人工智能学院; 中国科学院信息工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RSD-Poker提出结构自适应与移位鲁棒的残差策略认证框架,通过冻结残差族、学习划分及鲁棒映射,在不完全信息博弈中实现风险与效用双重证书,显著降低留出违规并提升弱效用。

AI 中文摘要

残差策略自适应提供了一种轻量级修改强参考策略的方式,但共享的尺度与固定的子群划分可能掩盖异质性退化,并在信息状态的部署混合发生变化时变得脆弱。我们提出RSD-Poker,一个结构自适应且移位鲁棒的认证框架,它冻结一组残差族与尺度,在独立的结构划分上学习策略可见的划分,并在校准标签加入前冻结该划分。每个候选-组对获得加权的同时上证书(针对锚定相对风险)与下证书(针对弱响应效用)。随后,在预先声明的部署组比例不确定性集上选择鲁棒的组到候选映射。在从每个冻结组的分布独立抽取的校准单元、校准前固定的候选库与划分以及组内条件不变性的条件下,所选映射以至少$1-\zeta_{risk}-\zeta_{util}$的概率满足其声明的混合鲁棒风险预算与效用证书。信息契约支持教师支持的变换与无教师的仅观测学生。保留的确定性24状态审计仍是精确回放诊断:经验零选择$\alpha=0.08$,将弱响应代理从4.2082提升至4.2889,且0/12留出阈值交叉。在分层留出状态上,学习划分的双重选择器将弱效用从全局双重认证下的4.4074提升至4.4936,并将留出违规从0.0215降至0.0078;其混合鲁棒变体达到违规0.0059。在五个仅观测检查点上,风险校准残差达到弱效用$4.3659\pm0.0177$与违规率$0.0178\pm0.0057$。

英文摘要

Residual policy adaptation provides a lightweight way to modify a strong reference policy, but a shared scale and a fixed subgroup partition can hide heterogeneous degradation and become fragile when the deployment mixture of information states changes. We introduce RSD-Poker, a structure-adaptive and shift-robust certification framework that freezes a bank of residual families and scales, learns a policy-visible partition on an independent structure split, and freezes that partition before calibration labels are joined. Each candidate-group pair receives a weighted simultaneous upper certificate for anchor-relative risk and a lower certificate for weak-response utility. A robust group-to-candidate map is then selected over a predeclared uncertainty set of deployment group proportions. Under independent calibration units drawn from each frozen group's law, a candidate bank and partition fixed before calibration, and invariant within-group conditionals, the selected map satisfies its declared mixture-robust risk budget and utility certificate with probability at least $1-ζ_{risk}-ζ_{util}$. The information contract supports both a teacher-backed transform and a teacher-free observation-only student. The retained deterministic 24-state audit remains an exact replay diagnostic: empirical-zero selects $α=0.08$, raising the weak-response proxy from 4.2082 to 4.2889 with $0/12$ held-out threshold crossings. On stratified held-out states, the learned-partition dual selector raises weak utility from 4.4074 under global dual certification to 4.4936 and lowers held-out violation from 0.0215 to 0.0078; its mixture-robust variant reaches violation 0.0059. Across five observation-only checkpoints, risk-calibrated residuals attain weak utility $4.3659\pm0.0177$ and violation rate $0.0178\pm0.0057$.

Comments36 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑