arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09907cs.LGmath.OC

Meta-LinEXP3:对抗性线性上下文赌博机的在线内在线学习

Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits

Hao Li, Jie Xu, Zheng Xie

首次发表
浏览论文内容

中文总结 AI 辅助

针对对抗性线性上下文赌博机,提出在线内在线元学习算法 Meta-LinEXP3,利用任务先验指导内部学习器,实现次线性遗憾并验证于高光谱采样。

中文摘要 AI 辅助

元学习已成为跨序列赌博机任务传递知识的有效范式。尽管在随机赌博机和非上下文对抗性赌博机方面已取得重大进展,但针对具有随机动作集的对抗性线性上下文赌博机(ALCBs)的元学习仍 largely 未被探索。为解决此问题,我们提出 Meta-LinEXP3,一种在线内在线算法,该算法从已完成任务中构建可预测的任务级先验,以指导内部 LinEXP3 学习器。对于已知上下文分布,我们开发了一种以策略为中心的估计器,实现了内在维度 $\nmathcal{O}(\sqrt{n})$ 的每任务遗憾界。对于未知分布,我们引入了一种仅使用过去数据的正则化矩估计器,其主导遗憾项为 $\mathcal{O}(n^{2/3})$,并具有明确的有限样本误差。我们进一步建立了先验准确性与迁移遗憾之间的直接联系,表明越来越准确的先验可跨任务产生次线性的迁移相关遗憾。实验证明了 Meta-LinEXP3 的有效性,包括其在结构化高光谱张量采样中的应用。

英文摘要

Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action sets remains largely unexplored. To address this problem, we propose Meta-LinEXP3, an online-within-online algorithm that constructs a predictable task-level prior from completed tasks to guide the inner LinEXP3 learner. For known context distributions, we develop a policy-centered estimator that achieves an intrinsic-dimension $\mathcal{O}(\sqrt{n})$ per-task regret bound. For unknown distributions, we introduce a past-only regularized moment estimator with an $\mathcal{O}(n^{2/3})$ leading regret term and explicit finite-sample error. We further establish a direct connection between prior accuracy and transfer regret, showing that increasingly accurate priors yield sublinear transfer-dependent regret across tasks. Experiments demonstrate the effectiveness of Meta-LinEXP3, including its application to structured hyperspectral tensor sampling.

发表机构

  • College of Science, National University of Defense Technology(国防科技大学理学院)

机构由 AI 辅助整理,请以论文原文为准。

↑