arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效在线词典序广义低秩矩阵老虎机

Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits

Bo Xue, Ji Cheng, Haodong Jing, Hongzong Li, Shuang Qiu

arXiv 2608.04324首次发表:更新:

发表机构

City University of Hong Kong; Northwestern Polytechnical University(香港城市大学; 西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对多优先目标的广义低秩矩阵老虎机问题,提出高效在线算法\textsc{Lexi-LowGLM},将估计器更新复杂度从$O(T^2)$降至$O(T)$,并通过实验验证其有效性与效率。

AI 中文摘要

本文研究具有多个优先目标的广义低秩矩阵老虎机。在每一轮中,学习者选择一个矩阵值臂并观察一个向量值奖励,其分量对应于具有不同优先级的多个目标。每个目标由特定于目标的广义低秩矩阵模型控制,学习者根据词典序偏好顺序评估臂,在较低级别目标之前优先考虑较高级别目标。我们提出了\textsc{Lexi-LowGLM},这是一种高效的在线算法,该算法首先估计特定于目标的低秩子空间,然后在降维特征空间中执行词典序学习。与现有单目标算法(使用所有历史观测值反复求解批量广义线性估计器)不同,\textsc{Lexi-LowGLM}通过在线牛顿步更新每个特定于目标的估计器,将T轮内的估计器更新复杂度从$O(T^2)$降低到$O(T)$。我们为每个目标$i\in[m]$建立了$\widetilde O\left(W_i^{\rm lex}\sqrt{m}\\,(d_1+d_2)r\sqrt{T}\right)$的后悔界,其中$r$是特定于目标的参数矩阵秩的上界,$W_i^{\rm lex}$表征词典序权衡效应。该界依赖于有效低秩维度$(d_1+d_2)r$而非环境维度$d_1d_2$。数值实验进一步验证了所提方法的有效性和计算效率。

英文摘要

This paper studies generalized low-rank matrix bandits with multiple prioritized objectives. At each round, the learner selects a matrix-valued arm and observes a vector-valued reward, whose components correspond to multiple objectives with different priority levels. Each objective is governed by an objective-specific generalized low-rank matrix model, and the learner evaluates arms according to a lexicographic preference order, prioritizing higher-level objectives before lower-level ones. We propose \textsc{Lexi-LowGLM}, an efficient online algorithm that first estimates objective-specific low-rank subspaces and then performs lexicographic learning in the reduced feature spaces. Unlike existing single-objective algorithms that repeatedly solve a batch generalized linear estimator using all historical observations, \textsc{Lexi-LowGLM} updates each objective-specific estimator via an online Newton step, reducing the estimator-update complexity over $T$ rounds from $O(T^2)$ to $O(T)$. We establish a regret bound of $\widetilde O\left(W_i^{\rm lex}\sqrt{m}\,(d_1+d_2)r\sqrt{T}\right)$ for each objective $i\in[m]$, where $r$ is an upper bound on the ranks of the objective-specific parameter matrices and $W_i^{\rm lex}$ characterizes the lexicographic trade-off effect. This bound depends on the effective low-rank dimension $(d_1+d_2)r$ rather than the ambient dimension $d_1d_2$. Numerical experiments further validate the effectiveness and computational efficiency of the proposed method.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑