在线逆线性优化的紧致遗憾界:基于多尺度矩阵权重
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
- CyberAgent(赛博艾坚特)
- National Institute of Informatics(国立信息学研究所)
- RIKEN(理化学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对固定未知线性效用的在线逆线性优化问题,提出基于多尺度矩阵权重的随机算法,实现期望遗憾 $O(\sqrt d)$,在常数因子内最优,且无需预知时间范围。
AI中文摘要:
我们研究具有固定未知线性效用的在线逆线性优化问题:在每一轮中,环境呈现一个紧致动作集,学习器从中推荐一个动作,然后环境返回一个在该集合上最大化效用的动作。当效用向量和动作位于 $d$ 维欧几里得单位球内时,我们给出一个随机化算法,其遗憾——相对于最优动作的累积效用缺口——在任意时间范围内期望为 $O(\sqrt d)$,且无需知道时间范围。根据已知的 $\Omega(\sqrt d)$ 下界(对于 $T\ge d$ 的时间范围),对 $d$ 的依赖在常数因子意义下是最优的。我们的算法在几何间隔的尺度上,对多项式特征空间维护矩阵乘法权重。它通过求解一个线性规划来选择推荐分布,并通过比较可用动作与反馈动作来更新其得分矩阵。在有理数预言机输出和反馈动作的情况下,相对于线性优化预言机可计算的实现保持了 $O(\sqrt d)$ 的遗憾界。是否能在维度、时间范围和输入长度的多项式时间内达到相同速率仍然是一个开放问题。
英文摘要:
We study online inverse linear optimization with a fixed unknown linear utility: in each round, an environment presents a compact action set, the learner recommends an action from it, and the environment returns an action that maximizes the utility over the same set. When the utility vector and the actions lie in the $d$-dimensional Euclidean unit ball, we give a randomized algorithm whose regret---the cumulative utility shortfall relative to optimal actions---is $O(\sqrt d)$ in expectation for every time horizon, without knowledge of the horizon. The dependence on $d$ is optimal up to a constant factor by the known $Ω(\sqrt d)$ lower bound for horizons $T\ge d$. Our algorithm maintains matrix multiplicative weights on polynomial feature spaces at geometrically spaced scales. It selects a recommendation distribution by solving a linear program and updates its score matrices by comparing the available actions with the feedback action. With rational oracle outputs and feedback actions, an implementation computable relative to a linear-optimization oracle preserves the $O(\sqrt d)$ regret bound. Whether the same rate is attainable with running time polynomial in the dimension, horizon, and input length remains open.