arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27910cs.AIcs.CLcs.GT

基于博弈论视角的AI对齐:一项综述

AI Alignment through a Game-theoretic Lens: A Survey

  • James Cook University(詹姆斯库克大学)
  • Western Sydney University(西悉尼大学)
  • Adelaide University(阿德莱德大学)
  • CSIRO(联邦科学与工业研究组织)
  • The University of Osaka(大阪大学)
  • Nanjing University of Science and Technology(南京理工大学)
  • Macquarie University(麦考瑞大学)
  • La Trobe University(拉筹伯大学)

机构由 AI 辅助整理,请以论文原文为准。

Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang

中文总结 AI 辅助

本综述从博弈论视角梳理AI对齐的近期进展,针对偏好多样性等三大挑战整合相关文献,明确了博弈论在AI对齐中的作用、不足及未来构建稳健AI系统的挑战。

中文摘要 AI 辅助

随着大型语言模型和能力日益增强的AI智能体被部署到高风险场景中,使其与复杂的人类价值观对齐已成为核心挑战。现有的对齐方法虽能有效提升有用性、无害性和可控性,但往往难以捕捉依赖于上下文、非传递性且受多方动态互动影响的现实偏好。本综述从博弈论视角审视AI对齐,具体围绕关键博弈论要素梳理近期进展,并从偏好多样性、对齐优先级和时间动态性这三个挑战维度整合相关文献。该视角明确了现有对齐方法在哪些方面真正得益于博弈论分析、该框架在哪些方面较为松散,以及在构建稳健、自适应且可验证的AI系统方面仍存在哪些挑战。

英文摘要

As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability, often struggle to capture real-world preferences that are context-dependent, non-transitive, and shaped by dynamic multi-party interactions. This survey reviews AI alignment through a game-theoretic lens. Specifically, it organizes recent progress around key game-theoretic elements and synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics. This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.

补充信息

↑