发表机构
IRIF, Université Paris Cité; Stellenbosch University; Max Planck Institute for Software Systems; School of Informatics, University of Edinburgh; Department of Computer Science, Oxford University(巴黎西岱大学; 斯坦伦布什大学; 马克斯·普朗克软件系统研究所; 爱丁堡大学信息学院; 牛津大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明偿付能力博弈的最优策略通常不必是最终周期的,否定了2012年的猜想,并给出了特定增益范围下的周期性、唯一性及可计算性结果。
AI 中文摘要
偿付能力博弈是无限状态马尔可夫决策过程上的一种赌博问题,其中状态 $n \in \mathbb{N}$ 代表投资者的财富。在每一轮中,投资者从有限动作集中选择一个动作,每个动作在区间 $\{-\ell,\ldots,m\}$ 内产生一个整数增益的分布。风险厌恶的投资者希望最小化最终破产(财富达到 $\le 0$)的概率。文献[Berger et al.]已证明存在无记忆确定性最优策略,但这些策略通常不是最终恒定的。即使在增益属于 $\{-2,\ldots,1\}$ 的特殊情况下,最优策略也可能需要在任意高的财富水平下使用两个不同的动作。我们证明偿付能力博弈中的最优策略通常不必是最终周期的(从而否定了Kučera在2012年的一个猜想)。在增益属于 $\{-3,\ldots,1\}$ 的情况下,最优策略可能是唯一的但非周期性的。对于增益属于 $\{-2,\ldots,1\}$ 的情况,总是存在一个最终周期的最优策略,其尾部是恒定的或在两个动作之间交替。最后,我们证明如果最优策略是唯一的,则它是可计算的。此外,对于任何 $\ell \in \mathbb{N}$,在增益属于 $\{-\ell,\ldots,1\}$ 的情况下,总是可以计算出(某个)最优策略。然而,一般情况下的可计算性仍然是一个开放问题。
英文摘要
Solvency games are a gambling problem on infinite-state Markov decision processes in which the state $n \in \mathbb{N}$ represents an investor's fortune. In every round, the investor chooses an action from a finite action set, and every action yields a distribution over integer-valued gains in an interval $\{-\ell,\ldots,m\}$. The risk-averse investor wants to minimise the probability of eventual ruin (reaching a fortune $\le 0$). It was shown in [Berger et al.] that memoryless deterministic optimal strategies exist, but they are not eventually constant in general. Even in the special case of gains in $\{-2,\ldots,1\}$, the optimal strategy may need to make use of two different actions at arbitrarily high fortunes. We show that optimal strategies in solvency games need not be ultimately periodic in general (thus disproving a 2012 conjecture of Kučera). Already in the case of gains in $\{-3,\ldots,1\}$, it is possible for the optimal strategy to be unique but aperiodic. For gains in $\{-2,\ldots,1\}$, there always exists an ultimately periodic optimal strategy whose tail is constant or alternates between two actions. Finally, we show that the optimal strategy is computable if it is unique. Moreover, (some) optimal strategy can always be computed in the case of gains in $\{-\ell,\ldots,1\}$ for any $\ell \in \mathbb{N}$. Computability in the general case however remains open.