arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23265cs.LG

重复先知不等式的无遗憾最优学习

Optimal No-Regret Learning for Repeated Prophet Inequality

Kun Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对前缀反馈下的重复先知不等式问题,提出一种结合经验逆向归纳与相对下降聚合的高效算法,实现与下界匹配的 $\tilde O(\sqrt T)$ 遗憾,并消除对盒子数量的多项式依赖。

中文摘要 AI 辅助

我们研究了在前缀反馈下的重复先知不等式问题。在总共 $T$ 轮中的每一轮,学习者会遇到从 $n$ 个盒子中独立抽取的新值,这些值来自未知的、支撑在 $[0,1]$ 上的分布,并按照固定顺序呈现,学习者必须不可撤销地接受其中一个值,且只能观察到直到其停止盒子为止的前缀。遗憾是针对知道分布的最优停止策略来衡量的。我们给出了一种高效算法,实现了 $\tilde O(\sqrt T)$ 的期望遗憾,与下界在对数因子内匹配。我们的算法通过接近最优的策略直接进行探索,将经验逆向归纳与盒子特定的到达奖励相结合。然后,一种相对下降聚合规则利用观察到的前缀的嵌套结构来保持探索,从而消除了对盒子数量 $n$ 的多项式依赖。这解决了 Liu 等人(2025)提出的一个开放问题。

英文摘要

We study repeated prophet inequalities under prefix feedback. In each of $T$ rounds, a learner encounters fresh values drawn independently from $n$ boxes with unknown $[0,1]$-supported distributions in a fixed order and must irrevocably accept one, observing only the prefix up to its stopping box. Regret is measured against the optimal stopping policy that knows the distributions. We give an efficient algorithm achieving $\widetilde O(\sqrt{T})$ expected regret, matching the lower bound up to logarithmic factors. Our algorithm explores directly through near-optimal policies, combining empirical backward induction with box-specific reach bonuses. A relative-drop aggregation rule then exploits the nesting structure of observed prefixes to preserve exploration, thereby removing the polynomial dependence on the box number $n$. This resolves an open question posed by Liu et al. (2025).

发表机构

  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑