五位专家预测与几何停止:概率构造与分析验证
Prediction with Five Experts and Geometric Stopping: A Probabilistic Construction and Analytic Verification
- University of Michigan(密歇根大学)
- University of Cyprus(塞浦路斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文用随机微积分和偏微分方程求解几何停止下五位专家预测的极限问题,给出极小极大遗憾渐近公式,并验证哈密顿-雅可比-贝尔曼方程及全局正则性。
AI中文摘要:
我们基于随机微积分和偏微分方程,提出了具有几何停止的极限五位专家预测问题的解。若 $\delta$ 为停止概率,则从平局初始分数出发的极小极大期望遗憾为 $45\pi^2/(512\sqrt{2\delta})+o(\delta^{-1/2})$(当 $\delta\downarrow0$ 时)。选择最佳和第三佳专家的对抗性秩控制在每个状态下最大化极限哈密顿量。遵循 Bayraktar、Ekren 和 Zhang(2020)的四位专家构造,我们将价值修正表示为退化斜反射布朗运动的贴现边界局部时间期望。控制其边界迹的双曲系统被推导并显式求解。为验证非线性 Hamilton--Jacobi--Bellman 方程,我们结合等式方向、双曲比较原理以及跨控制的平均化。这些论证将 48 个区域控制不等式降维为低维边界问题,其余符号由显式单调性论证确定。我们证明了跨区域界面和秩排序变化的全局 $C^2$ 正则性。我们还表明,当两个最高分数相等且第三和第四高分数相等时,交替 COMB 控制恰好达到极限哈密顿量最大值。
英文摘要:
We present a solution of the limiting five-expert prediction problem with geometric stopping, based on stochastic calculus and partial differential equations. If $δ$ is the stopping probability, the minimax expected regret from tied initial scores is $45π^2/(512\sqrt{2δ})+o(δ^{-1/2})$ as $δ\downarrow0$. The adversarial rank control selecting the best and third-best experts maximizes the limiting Hamiltonian at every state. Following the four-expert construction of Bayraktar, Ekren, and Zhang (2020), we represent the value correction as a discounted boundary-local-time expectation for a degenerate obliquely reflected Brownian motion. The hyperbolic systems governing its boundary traces are derived and solved explicitly. To verify the nonlinear Hamilton--Jacobi--Bellman equation, we combine heat-equation minimum principles in equality directions, ODE minimum principles applied to convex combinations of multiple controls, and comparisons across controls. These arguments reduce the 48 regional control inequalities to a few lower-dimensional boundary problems, with the remaining signs established by explicit monotonicity arguments and by factorization of multivariate integral kernels that are positive polynomials in some of their variables. We prove global $C^2$ regularity across both regional interfaces and changes in rank ordering. We also show that the alternating COMB control attains the limiting Hamiltonian maximum exactly when the two highest scores coincide and the third- and fourth-highest scores coincide.