AI 中文总结
本研究提出T-ESTOR和T-BSTOR算法,解决低秩矩阵与张量老虎机中未知奖励链接的问题,分别实现最优遗憾界,并通过实验验证了结构化估计的优势。
AI 中文摘要
低秩矩阵和张量老虎机利用结构化交互,但通常假设奖励链接已知。近期的单指标老虎机方法能适应未知链接,却不直接利用矩阵或张量秩。我们通过研究在已知正则候选分布和有限方差噪声下,具有未知共享Lipschitz链接和低秩指标参数的随机矩阵和张量老虎机来填补这一空白。对于单调链接,T-ESTOR结合了鲁棒的秩自适应Stein估计与基于epoch的贪心选择。在精确选择分数访问和一致正的选择设计Stein信号下,它实现了平方根遗憾,其维度依赖性由低秩结构决定。对于每个可接受的设计,在大时间范围下,单调下界在秩、维度和时间范围依赖性上与对数因子匹配,对于固定菜单大小和模型/设计常数。对于非单调链接,在非零基础律Stein信号下,T-BSTOR结合了结构化估计与鲁棒的基于分箱的学习,并在固定维度、菜单大小和模型/设计常数下达到最优的$\widetilde{O}(T^{2/3})$时间范围速率。合成和基于CCLE的实验展示了结构化估计相对于向量化和竞争性单指标基线方法的优势。
英文摘要
Low-rank matrix and tensor bandits exploit structured interactions but typically assume a known reward link. Recent single-index bandit methods accommodate unknown links without directly exploiting matrix or tensor rank. We address this gap by studying stochastic matrix and tensor bandits with an unknown shared Lipschitz link and a low-rank index parameter under known regular candidate distributions and finite-variance noise. For monotone links, T-ESTOR combines robust, rank-adaptive Stein estimation with epoch-based greedy selection. Under exact selected-score access and a uniformly positive selected-design Stein signal, it achieves square-root regret with dimension dependence determined by the low-rank structure. For every admissible design, the monotone lower bound matches the rank, dimension, and horizon dependence up to logarithmic factors at large horizons, for fixed menu size and model/design constants. For nonmonotone links under a nonzero base-law Stein signal, T-BSTOR combines structured estimation with robust bin-based learning and attains the optimal $\widetilde{O}(T^{2/3})$ horizon rate for fixed dimensions, menu size, and model/design constants. Synthetic and CCLE-based experiments illustrate the benefits of structured estimation relative to vectorized and competing single-index baseline methods.