arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14761cs.GTcs.LG

无偏性约束下的CFR:持久公共机会调度的确定性保证

CFR without Unbiasedness: Deterministic Guarantees for Persistent Public-Chance Schedules

Jiaxing Guo, Lei Ye

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对持久公共机会调度建立CFR的确定性保证,提出目标转移定理,经实验验证其调度在扑克残局中优于重新洗牌,为设计审计持久CFR提供基础。

中文摘要 AI 辅助

在有限的公共机会切口处,反事实遗憾最小化(CFR)必须在每次遗憾更新前选择要评估的结果数量。精确评估在单个策略剖面处理完整切口;持久部分评估则在不断演化的剖面间处理固定的无放回顺序,每个周期覆盖所有结果一次,但其反馈通常存在条件偏差,因为早期批次会影响后期批次所见的剖面。我们针对均匀、非嵌套的加性公共切口建立确定性目标转移定理,该定理通过所传递反馈的遗憾以及将前缀覆盖差异与实现策略路径上的运动耦合的公共借方项,来界定完整切口的可利用性。因此,在预定平均权重下,连续平衡调度对加性符号遗憾匹配(RM)和RM+收敛,而固定RM+构造证明了差异-路径乘积在一般情况下是必要的。该定理的分量解析形式可将执行轨迹转换为数值可利用性证书。在两个已发布的一对一无限注德州扑克转牌残局中,尽管周期覆盖相同,持久顺序相比重新洗牌有显著提升,且在每个注册的浅层匹配预算比较中,部分覆盖均获胜。深度研究发现32到64个完整切口结果预算之间存在交叉点,超过该点后完整覆盖占优。这些结果将公共机会的宽度和顺序表征为学习变量,并为设计和审计持久CFR调度提供确定性基础。

英文摘要

At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update. Exact evaluation processes the full cut at one strategy profile; persistent partial evaluation processes a fixed without-replacement order across evolving profiles. The latter covers every outcome once per epoch, yet its feedback is generally conditionally biased because earlier batches influence the profiles seen by later batches. We establish a deterministic target-transfer theorem for uniform, nonnested additive public cuts. The theorem bounds full-cut exploitability by regret on the delivered feedback and a public-debit term that couples prefix coverage discrepancy with motion along the realized strategy path. Consecutively balanced schedules consequently converge for additive signed regret matching (RM) and RM+ under predetermined averaging weights, while a fixed RM+ construction proves that the discrepancy--path product is necessary in general. A component-resolved form of the theorem converts an execution trace into a numerical exploitability certificate. On two released heads-up no-limit hold'em turn endgames, persistent order improves substantially over fresh reshuffling despite identical epochwise coverage, and partial coverage wins every registered shallow matched-budget comparison. A depth study locates a crossover between 32 and 64 full-cut outcome budgets, after which complete coverage dominates. These results characterize public-chance width and order as learning variables and provide a deterministic basis for designing and auditing persistent CFR schedules.

发表机构

  • Imperial College London(伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑