关于鲁棒多臂老虎机计算可处理性的研究
On the Computational Tractability of Robust Bandits
浏览论文内容
中文总结 AI 辅助
本文研究鲁棒多臂老虎机问题的计算可处理性,发现一个特例存在多项式时间且遗憾为Õ(√T)的学习器,而其若干推广为NP难,为AI对齐提供计算高效学习器迈出一步。
中文摘要 AI 辅助
当环境不属于学习者的假设类别时,学习问题通常通过不可知学习保证来处理。然而,对于监督学习之外的任何问题,不可知保证都难以获得。最近,不精确多臂老虎机(Kosoy,2025)(后在Appel和Kosoy,2025中更名为鲁棒多臂老虎机)被引入作为多臂老虎机设置中不可实现学习的另一种方法,并针对一大类问题展示了Θ(√T)遗憾的学习器。然而,未提供计算保证。在本文中,我们确定了一个特例,该特例允许具有Õ(√T)遗憾的多项式时间学习器。我们还表明,该特例的几个小规模推广是NP难的,从而表明该特例处于可处理性的边界。最近有人提出(Kosoy,2018),针对不可实现学习问题的计算高效学习器对于解决AI对齐问题至关重要。这项工作正是朝着该方向迈出的一小步。
英文摘要
Learning when the environment does not belong to the learner's hypothesis class is typically handled using agnostic learning guarantees. However, for anything beyond supervised learning, agnostic guarantees are difficult to come by. Recently, imprecise bandits (Kosoy, 2025) (later renamed to robust bandits in Appel and Kosoy, 2025) were introduced as another approach to unrealizable learning in the bandits setting and a $Θ(\sqrt{T})$ regret learner was shown for a large class. However, no computational guarantees were provided. In this paper we identify a special case that admits a polynomial-time learner with $\tilde{O}(\sqrt{T})$ regret. We also show that several small generalizations of this special case are NP-hard thus indicating that the special case is at the boundary of what is tractable. It has been recently suggested (Kosoy, 2018) that computationally efficient learners for unrealizable learning problems are crucial for solving the AI alignment problem. This work is a small step in that direction.
发表机构
- Technion(以色列理工学院)
- CORAL(计算理性智能体实验室)
机构由 AI 辅助整理,请以论文原文为准。