arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08735cs.LGcs.DS

最优且高效的在线逆向优化

Optimal and Efficient Online Inverse Optimization

Anupam Gupta, Guru Guruganesh, Honghao Lin, Vahab Mirrokni, Renato Paes Leme, David P. Woodruff

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种确定性多项式时间算法,在在线逆向线性优化中达到最优遗憾 $O(\sqrt d)$,通过变度量更新撤销机制解决了Sakaue提出的开放问题。

中文摘要 AI 辅助

在在线逆向线性优化中,学习器推荐一个动作,然后观察一位专家的选择,该专家在 $\mathbb{R}^{d}$ 上最大化一个固定但未知的线性目标;目标是在不直接观察该目标的情况下学会优化它。Sakaue 最近通过一种每轮进行 $(dT)^{O(d)}$ 次线性优化的随机算法获得了最优遗憾 $O(\sqrt d)$,并提出了能否在多项式时间内达到该遗憾的问题。我们给出了肯定回答:我们的确定性算法对每个时间范围 $T$ 都达到遗憾 $O(\sqrt d)$,并且运行时间在 $d$ 和 $T$ 上是多项式的。它是 Sakaue 等人和 Cai 等人的变度量算法的一个变体,其中当查询点移动得离更新位置足够远时,度量更新会被撤销。

英文摘要

In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on $\mathbb{R}^{d}$; the goal is to learn to optimize this objective without observing it. Sakaue recently obtained the optimal regret $O(\sqrt d)$ with a randomized algorithm making $(dT)^{O(d)}$ linear optimizations per round, and asked whether it can be attained in polynomial time. We answer positively: our deterministic algorithm has regret $O(\sqrt d)$ for every horizon $T$ and runs in time polynomial in $d$ and $T$. It is a variant of the variable-metric algorithms of Sakaue et al.\ and Cai et al., in which a metric update is revoked once the query point moves far enough from where the update was made.

发表机构

  • Google Research(谷歌研究院)
  • New York University(纽约大学)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑