arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28881cs.AI

不完美对齐下价值的脆弱性

Fragility of Value under Imperfect Alignment

Winter Cross, Léo Cymbalista, Alfred Harwood, Jose Faustino

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对AI系统与人类价值对齐问题,建立模型明确人类价值函数与代理条件准确度的相关条件,凸显过度优化风险,提出采用quantilizers等限制优化压力的AI设计方案。

中文摘要 AI 辅助

随着AI系统承担越来越多的责任,确保这些系统与人类对齐变得愈发重要。AI安全领域普遍存在一种担忧:人类价值是脆弱的——即过度优化人类价值的不完美代理会导致灾难性结果。本文提出了一种对齐问题模型,其中智能体经过理想化的对齐训练,保证其价值函数在优化世界前满足代理条件。主要结果明确了人类价值函数的条件,以及若干代理条件的准确度,在此条件下,具有η灾难性价值函数(该函数保证在优化能力极限下将人类价值的期望降至η以下)的智能体将被部署。研究结果凸显了过度优化的风险,推动AI设计需限制优化压力,例如采用quantilizers(分位数器),而非仅依赖部署前训练。

英文摘要

As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $η$-catastrophic value function, one that is guaranteed to take the expectation of human value below $η$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.

补充信息

↑