有限行动时间智能体的AI安全考量
AI Safety Considerations for Agents With Limited Time to Act
浏览论文内容
中文总结 AI 辅助
本文讨论在部分可观测且行动时间有限的环境中,智能体无关的安全保证的理论界限,通过两个场景证明即使完美智能体也无法保证安全,强调环境与安全行动需与智能体共同考虑。
中文摘要 AI 辅助
随着关于AI对齐的公开讨论日益增多,近期工作试图提出能够安全行为的具体AI架构。然而,那些看似证明对齐的论点大多忽略了智能体需要行动的环境。我们讨论了在只能部分观察且需要在有限时间内采取行动的环境中,与智能体无关的安全保证的理论界限。我们引入了两个现实场景,一个具有无限状态空间,另一个具有信号混合。在这些场景中,我们证明即使是一个完美的智能体也无法保证安全行为。我们将论证,对于任何AI安全或对齐的证明,环境和相关的安全行动都需要与智能体一起被具体考虑。
英文摘要
In the wake of the increasingly public discussion about AI alignment, recent work has tried to propose specific AI architectures that behave safely. However, the proposed arguments that seemingly demonstrate proved alignment mostly neglect the environment the agent needs to act in. We discuss theoretical bounds for agent-agnostic safety guarantees in environments that can only be partially observed and within which an action is required within limited time. We introduce two realistic scenarios, one with an infinite state space and one with signal mixture. In these scenarios, we prove that even a perfect agent cannot guarantee safe behaviour. It will be argued that for any proof of AI safety or alignment, the environment and associated safe actions need to be specifically considered together with the agent.
发表机构
- The Alan Turing Institute(艾伦·图灵研究所)
机构由 AI 辅助整理,请以论文原文为准。