arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35622stat.MLcs.LG

单指标老虎机中的诱发与决策几何

Elicitation and Decision Geometry in Single-Index Bandits

Sakshi Arya, Cheng Soon Ong

首次发表
浏览论文内容

中文总结 AI 辅助

针对双臂单指标上下文老虎机,提出自然边界学习(NBL)方法,利用序贯斯坦对比直接学习最优边界,在决策稳定性条件下实现对数遗憾,并通过数值实验验证。

中文摘要 AI 辅助

我们研究具有臂特定单指标和共享未知单调链接的双臂上下文老虎机。单调性使得最优行动仅取决于指标方向之间的对比,因此无需估计臂特定奖励函数。我们引入自然边界学习(NBL),一种贪心程序,使用序贯斯坦对比直接学习最优边界,而无需估计奖励函数或共同链接。我们通过决策稳定性系数刻画NBL的局部黎曼动力学,该系数平衡臂分离、链接几何和上下文分布。我们表明这种稳定性与底层凸势的诱发几何相关。在局部决策稳定性下,NBL向最优边界收缩并实现$O(\log n)$期望遗憾。数值实验说明了预测的稳定性机制,并在链接误设下将NBL与参数化贪心基准进行比较。

英文摘要

We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (NBL), a greedy procedure that uses a sequential Stein contrast to learn the optimal boundary directly, without estimating the reward functions or the common link. We characterize the local Riemannian dynamics of NBL through a decision stability coefficient balancing arm separation, link geometry, and the context distribution. We show that this stability is connected to the elicitation geometry of the underlying convex potential. Under local decision stability, NBL contracts toward the optimal boundary and achieves $O(\log n)$ expected regret. Numerical experiments illustrate the predicted stability regimes and compare NBL with a parametric greedy benchmark under link misspecification.

发表机构

  • Case Western Reserve University(凯斯西储大学)
  • CSIRO(澳大利亚联邦科学与工业研究组织)

机构由 AI 辅助整理,请以论文原文为准。

↑