通过信息导向采样对齐AI智能体
Aligning AI Agents via Information-Directed Sampling
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出老虎机对齐问题框架,研究智能体在环境与人类偏好均未知时如何权衡探索与利用,并通过理论和实验证明信息导向采样优于朴素探索和Thompson采样。
AI中文摘要:
AI系统令人瞩目的成就使AI对齐这一议题受到关注:即将“超级智能”AI智能体的行动与人类利益对齐。现有对齐领域的许多框架和算法要么在短视时间范围内研究该问题,要么孤立地研究从人类反馈中学习,依赖智能体已完美识别环境这一人为假设。作为解决这些局限性的起点,我们定义了一类老虎机对齐问题,作为经典多臂老虎机问题的扩展。老虎机对齐问题涉及一个智能体,其任务是通过与环境及人类交互来最大化长期期望奖励,而环境和人类都包含智能体最初未知的细节与偏好。环境中行动的奖励取决于观测结果和人类偏好两者。此外,向人类查询以学习偏好会产生成本。因此,一个有效的智能体应当智能地权衡探索(对环境和对人类)与利用。我们在一个类似于beta-Bernoulli老虎机的玩具老虎机对齐问题中,从理论和实证两方面研究了这些权衡。我们证明,反映当前实践的朴素探索算法以及甚至被推崇的算法如Thompson采样都无法为该问题提供可接受的解决方案,而信息导向采样则取得了有利的遗憾。
英文摘要:
The staggering feats of AI systems have brought to attention the topic of AI Alignment: aligning a "superintelligent" AI agent's actions with humanity's interests. Many existing frameworks/algorithms in alignment study the problem on a myopic horizon or study learning from human feedback in isolation, relying on the contrived assumption that the agent has already perfectly identified the environment. As a starting point to address these limitations, we define a class of bandit alignment problems as an extension of classic multi-armed bandit problems. A bandit alignment problem involves an agent tasked with maximizing long-run expected reward by interacting with an environment and a human, both involving details/preferences initially unknown to the agent. The reward of actions in the environment depends on both observed outcomes and human preferences. Furthermore, costs are associated with querying the human to learn preferences. Therefore, an effective agent ought to intelligently trade-off exploration (of the environment and human) and exploitation. We study these trade-offs theoretically and empirically in a toy bandit alignment problem which resembles the beta-Bernoulli bandit. We demonstrate while naive exploration algorithms which reflect current practices and even touted algorithms such as Thompson sampling both fail to provide acceptable solutions to this problem, information-directed sampling achieves favorable regret.