arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

交互式对齐

Interactive Alignment

Sylvain Chassang

arXiv 2607.25019首次发表:更新:

发表机构

Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究交互式主体与人类福祉的长期对齐问题,采用人工智能主体模拟和进化博弈论框架两种方法,结果显示进化博弈论可近似宪法主体经济动态,实用规范执行能更有效维持长期对齐。

AI 中文摘要

本文研究了包括人工智能系统、团队、公司和政府在内的交互式主体与人类福祉的长期对齐问题。开发了一个农业游戏,主体群体进行种植、交易和扩张决策,要在向人类转移和自身扩张投资间分配最终产出,因向人类转移减少扩张资源,进化力量倾向于选择不利于对齐行为的结果。核心问题是能否设计主体关于分享和交易的宪法原则以使长期对齐持续。采用两种互补方法研究:一是开发人工智能主体模拟,主体偏好由书面宪法规定并由大语言模型解释;二是引入可处理的进化博弈论框架。结果表明进化博弈论能有效近似宪法主体经济动态,实用规范执行比简单利他主义或无条件利他执行更能有效维持长期对齐。

英文摘要

This paper is interested in the long-run alignment of populations of interactive agents (in particular AIs, but also teams, firms, and governments) with human welfare. Formally, it studies a farming game in which a population of agents make planting, trading, and expansion choices. The key alignment choice lies in how much final output to send to humans, and how much to invest in expansion. Because human welfare comes at the cost of expansion, this creates evolutionary pressure against alignment. The main question is whether it is possible to set up agents' constitutional principles regarding sharing and trading to ensure that alignment survives in the long run. The paper uses two complementary strategies to investigate the question: an AI-agent simulation where agents' preferences are described by a constitution and interpreted via an LLM; and a tractable analytical evolutionary game theory framework, allowing for rapid and intuitive exploration of the space of agent preferences. The analysis suggests that tools from evolutionary game theory provide a useful approximation of interactive agent economies, and that pragmatic norm enforcement shows promise in maintaining long-term alignment over simpler forms of altruism and altruistic enforcement.

Comments63 pages, 21 Figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑