arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于不断演变的社会规范的人工智能价值对齐

AI Value Alignment for Evolving Social Norms

Nenad Tomašev, Matija Franklin, Simon Osindero

arXiv 2607.18506首次发表:更新:

AI 中文总结

研究人工智能价值对齐对社会规范演变的影响,引入基于社会物理学的数学建模框架,通过解析与模拟分析长期后果,强调价值锁定等风险,倡导用此模型作认知桥梁助力社会技术预见及大规模智能体评估。

AI 中文摘要

人工智能对齐对于先进人工智能系统的安全部署至关重要。鉴于价值观和偏好会随时间、文化、社会角色和背景而变化,我们需要更好地理解人工智能对齐可能产生的长期后果,尤其是考虑到个性化人工智能助手未来可能的广泛使用。我们引入了一个灵活且可扩展的数学建模框架,该框架基于社会物理学,旨在回答在频繁使用人工智能的假设下,关于人类群体中不断演变的社会规范的宏观层面问题。我们的分析部分是解析的,部分是模拟的,这使我们能够在各种不同的初始假设下描述长期动态后果。我们强调了在非适应性对齐公式中显著存在的价值锁定和规范模式崩溃的风险。除了对齐之外,我们主张更广泛地采用这类社会物理学模型作为一种认知桥梁:为通用人工智能未来的社会技术预见实现快速、严格且基于定量的假设检验,并作为更具计算成本的大规模智能体评估的易于处理的先驱。

英文摘要

AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aimed at answering macro-level questions regarding the evolving social norms in human populations under the assumption of frequent AI use. Our analysis is part-analytical, and part-simulation, enabling us to characterize the long-term dynamical consequences under a diverse set of starting assumptions. We highlight the risk of value lock-in, and normative mode collapse, prominently featured in non-adaptive alignment formulations. Beyond alignment, we advocate for the wider adoption of these kinds of social physics models as an epistemic bridge: enabling rapid, rigorous, and quantitatively-grounded hypothesis testing for sociotechnical foresight in general AI futures, and acting as a tractable precursor to more computationally expensive large-scale agentic evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑