arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于对齐与控制的机制设计

Mechanism Design for Alignment and Control

Dirk Bergemann, Andrew Koh, Stephen Morris

arXiv 2609.01595首次发表:更新:

发表机构

Yale; Columbia; Google DeepMind; MIT(耶鲁大学; 哥伦比亚大学; 谷歌DeepMind; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种针对偏好与能力未知的AI智能体的机制设计框架,结合显示原理等理论,将其应用于怠惰、权衡、同伴评分等五类典型场景,为智能体对齐与控制提供了理论与方法支撑。

AI 中文摘要

我们开发了一种针对AI智能体的机制设计框架,这类智能体的对齐(偏好)与能力(可行行动及信息)未知。我们希望此类智能体代表我们行动,因此机制必须同时激励其诚实与服从。一种单侧模仿结构——能力可被隐瞒但无法伪造——产生了显示原理、通过嵌套循环单调性刻画可实施策略的方法,以及引出高阶信念可约束多个智能体的条件。我们将该框架应用于以下典型示例:(i)能力更强的智能体假装能力更弱的“怠惰(sandbagging)”;(ii)对齐与可解释性的权衡,二者在工具中为替代关系但在价值中为互补关系;(iii)通过同伴评分进行约束;(iv)耦合奖励以诱导多个智能体间的竞争;(v)可扩展监督与奖励塑造。

英文摘要

We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑