递归博弈近最优策略的单项式族的存在性与计算
Existence and computation of monomial families of near-optimal strategies for recursive games
AI总结:
本文为递归博弈近最优策略的单项式族正则性定理提供初等证明,并提出针对固定活跃状态数的有理博弈的确定性多项式时间算法,可精确计算该单项式族。
AI中文摘要:
在Everett定义的有限递归博弈中,对任意ε>0,双方均存在平稳ε-最优策略。Frederiksen与Miltersen强化了该结论,证明所有足够小的ε对应的策略可由有限个单项式编码:在每个状态下,除可能一个行动概率外,其余均为常数乘以ε的整数次幂,该有限符号对象可指定任意足够小精度下的策略,其证明使用半代数选择与Puiseux级数。本文为递归博弈的该正则性定理提供了一种初等替代证明:从保证收益向量通过Everett单侧区域趋近值的平稳策略出发,固定其支撑后,对每个纯平稳响应,将所有吸收概率表示为非负系数且有共同正分母的定向森林多项式之商,每个收益是这些商的固定带符号线性组合;随后将有限个森林单项式的渐近阶压缩为一个整数权重向量,该证明未使用半代数选择或Puiseux级数。此外,对于具有固定N个活跃状态的有理博弈,本文提出一种确定性多项式时间算法,可精确计算单项式族,返回所有代数系数的有序实单变量表示,其表示长度与运行时间最多为L^{(N+1)^{O(N)}},其中L为输入长度。
英文摘要:
In a finite recursive game in the sense of Everett, both players have stationary epsilon-optimal strategies for every epsilon>0. Frederiksen and Miltersen strengthened this result by showing that the strategies for all sufficiently small epsilon can be encoded by finitely many monomials: at every state, all but possibly one of the action probabilities are constants times integer powers of epsilon. The resulting finite symbolic object specifies a strategy for every sufficiently small accuracy. Their proof uses semialgebraic selection and Puiseux series. We give an alternative elementary proof of this regularity theorem for recursive games. We start with stationary strategies that guarantee vectors approaching the value through Everett's one-sided region. After fixing their support, we express, for each pure stationary reply, all absorption probabilities as quotients of directed-forest polynomials with nonnegative coefficients and a common positive denominator. Each payoff is a fixed signed linear combination of these quotients. We then compress the asymptotic orders of the finitely many forest monomials into one integer weight vector. This proof uses neither semialgebraic selection nor Puiseux series. Furthermore, for rational games with a fixed number N of active states, we present a deterministic polynomial-time algorithm that computes a monomial family exactly. It returns all algebraic coefficients in one ordered real univariate representation. The representation length and running time are at most L^{(N+1)^{O(N)}}, where L is the input length.