发表机构
Carnegie Mellon University; University of California Irvine(卡内基梅隆大学; 加州大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究现实世界中人类与人工智能交互的动态本质,主张从静态模拟对齐转向交互式互补对齐,通过对比现有对齐与轨迹级视图形式化差距,借鉴多学科见解指出挑战,最后概述开发与人类交互对齐的人工智能系统的研究议程。
AI 中文摘要
当前的对齐方法通常侧重于使用人类偏好的静态表示来模拟人类行为,未能捕捉现实世界中人类与人工智能交互的动态、上下文相关的本质。本文主张从静态和模拟对齐转向交互式和互补对齐,偏好通过交互产生,对齐不仅仅由满足偏好来定义。首先通过将现有对齐与轨迹级视图对比来形式化这一差距,其中人类和模型行为随时间共同演变。由于现有机器学习公式未充分捕捉这些交互动态,我们基于跨学科研讨会的见解来阐述这一观点。借鉴人类协作的社会科学描述,指出人机系统放大了这些动态,带来新的不对称性和协调挑战。最后概述了一个研究议程,以开发在交互中与人类对齐的人工智能系统,这需要机器学习与社会和决策科学的跨学科综合。
英文摘要
Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions. In this paper, we argue for a shift from static and emulative to interactive and complementary alignment, where preferences emerge through interaction and alignment is defined not by satisfying preferences alone. We first formalize this gap by contrasting existing alignment with a trajectory-level view in which human and model behavior co-evolve over time. Because these interaction dynamics have not been adequately captured within existing ML formulations, we ground this perspective in insights from an interdisciplinary workshop. We draw on lessons from social-science accounts of human-human collaboration and then argue that human-AI systems amplify these dynamics, introducing new asymmetries that make reasoning about uncertainty harder and introduce new coordination challenges. Based on these lessons and new challenges, we conclude by outlining a research agenda for developing AI systems that align with humans in interaction, requiring an interdisciplinary synthesis of machine learning and the social and decision sciences.