发表机构
Tilburg University(蒂尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对随机游走停止问题,提出基于样本的先知不等式,给出无限时域常数$(K/(K+1))^{K+1}$及有限时域结果,首次精确参数化样本数影响。
AI 中文摘要
我们研究了基于样本信息的随机游走奖励停止问题中的先知不等式。目标是在随机游走(具有独立同分布增量)中尽可能接近其最大值时停止,性能通过停止时的期望奖励与期望真实最大值之比来衡量。我们考虑一个基于样本的模型,其中增量分布未知,决策者可以访问奖励过程的$K$条独立样本路径。对于无限时域设置,我们建立了一个尖锐的先知不等式,其常数为$(K/(K+1))^{K+1}$。该保证通过基于随机游走阶梯高度分解的随机停止规则实现。当$K\to\infty$时,这恢复了全信息设置中的经典$1/e$先知不等式。对于有限时域设置,即过程在$n$步后终止,我们首先证明了一个紧的无信息先知不等式,其常数为$1/H_n$,其中$H_n$是第$n$个调和数。对于$K\ge1$个样本,我们证明先知常数$1/4$是可达到的。最后,我们证明,在$K$个样本下,对于$n\ge 2K^2$,先知常数至多为$(K/(K+1))^{K+1}+(6+6H_K)/H_n$,这意味着当$n\to\infty$时收敛到无限时域常数。我们的方法结合了随机游走理论(包括阶梯高度和Spitzer恒等式)与线性规划对偶性。我们的结果对随机游走停止理论以及基于样本的关联奖励先知不等式(这一领域在很大程度上仍未被探索)做出了贡献。据我们所知,我们紧的基于样本的先知不等式是第一个性能精确地由可用样本数量参数化(而非仅依赖于常数)的结果。
英文摘要
We study prophet inequalities for a random walk reward stopping problem with sample-based information. The goal is to stop as close as possible to the maximum of a random walk with i.i.d. increments, measuring performance by the ratio between the expected reward when stopping and the expected true maximum. We consider a sample-based model in which the increment distribution is unknown and the decision maker has access to $K$ independent sample paths of the reward process. For the infinite-horizon setting, we establish a sharp prophet inequality with constant $(K/(K+1))^{K+1}$. The guarantee is attained by a randomised stopping rule based on the ladder height decomposition of random walks. As $K\to\infty$, this recovers the classical $1/e$ prophet inequality from the full-information setting. For the finite-horizon setting, where the process terminates after $n$ steps, we first prove a tight no-information prophet inequality with constant $1/H_n$, where $H_n$ is the $n$-th harmonic number. For $K\ge1$ samples, we show that a prophet constant of $1/4$ is attainable. Finally, we prove that, with $K$ samples, the prophet constant is at most $(K/(K+1))^{K+1}+(6+6H_K)/H_n$ for $n\ge 2K^2$, implying convergence to the infinite-horizon constant as $n\to\infty$. Our approach combines random walk theory, including ladder heights and Spitzer's identity, with linear programming duality. Our results contribute to random walk stopping theory and to sample-based prophet inequalities for correlated rewards, an area that remains largely unexplored. To the best of our knowledge, our tight sample-based prophet inequalities are the first whose performance is parameterised exactly, rather than only up to constants, by the number of available samples.
CommentsAccepted at the 22nd Conference on Web and Internet Economics (WINE 2026)