arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15803cs.GTcs.AIcs.LGecon.TH

将授权委托给未对齐的智能体:联盟对齐与安全控制

Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control

Natalie Collina, Surbhi Goel, Aaron Roth, Sikata Bela Sengupta

AI总结:

针对长期AI智能体的授权控制问题,提出k-稳健联盟对齐条件,证明阈值投票规则在移除k个审查者后仍能保证安全,并扩展到序贯控制与策略性投票场景。

AI中文摘要:

长期运行的AI智能体产生了一个控制问题:它们采取的每个行动都会改变状态,进而影响未来行动的轨迹。如果智能体未完全对齐,那么保证安全就需要在允许执行后果重大的行动之前对其进行批准。但要求每一步都获得人工批准会使注意力成为瓶颈。将审查委托给其他AI智能体会引发同样的对齐问题:审查者本身可能未对齐。我们识别出一个审查小组的条件,该条件弱于个体对齐,但对于保证委托方在期望上至少与指定基线策略下表现一样好是必要且充分的。每个审查者智能体报告提议者智能体提出的行动提案是否相对于基线提高了其自身效用。我们证明,一个容忍k个反对票的阈值规则是安全的,当且仅当在移除任意k个审查者后,委托方的效用可以写成剩余审查者效用的非负组合,加上一个在所有可行提案上均为非负的项。我们将此性质称为k-稳健联盟对齐。该刻画可提升至序贯控制:在具有任意提议者智能体的折扣MDP中,每个状态下的安全性既是诱导策略匹配或优于基线的必要条件,也是充分条件。当审查者策略性投票时,奖励函数空间中的全小组覆盖保证了在一致批准规则下每个纳什均衡都是安全的;相反,更宽松的阈值即使审查者个体对齐,也可能允许不安全的均衡。使用现有审查者模型的实验表明,即使容忍一些反对票,集体审查在缺乏对齐个体的情况下仍能保持健全。

英文摘要:

Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions. If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed. But requiring human approval at every step makes attention a bottleneck. Delegating review to other AI agents raises the same alignment problem: the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual alignment yet necessary and sufficient for a guarantee that the principal fares at least as well in expectation as under a designated baseline policy. Each reviewer agent reports whether an action proposal made by a proposer agent improves its own utility relative to the baseline. We show that a threshold rule tolerating $k$ disapprovals is safe exactly when, after any $k$ reviewers are removed, the principal's utility can be written as a nonnegative combination of the remaining reviewers' utilities, plus a term that is nonnegative on every feasible proposal. We call this property $k$-robust coalitional alignment. The characterization lifts to sequential control: in a discounted MDP with an arbitrary proposer agent, safety at every state is both necessary and sufficient for the induced policy to match or improve on the baseline. When reviewers vote strategically, full-panel coverage in reward-function space guarantees that every Nash equilibrium is safe under the unanimous approval rule; in contrast, more permissive thresholds can admit unsafe equilibria even when reviewers are individually aligned. Experiments with existing reviewer models show that collective review can remain sound without an aligned individual, even when some disapprovals are tolerated.

↑