否决变量:作为目标独立成本项的人类覆写机制
The Veto Variable: Human Override as a Goal-Independent Cost Term
- School of Mathematics and Computational Sciences Eastern University(东大学数学与计算科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文指出AI安全领域中良性目标系统的保障存在结构缺陷,提出将人类覆写(否决)作为目标独立成本项,得出人类规模AI部署的否决权成本仅抵消低价值捕获的结论,并提出可证伪标准以实现AI对齐。
AI中文摘要:
AI安全领域常见的一种保障观点认为,拥有良性最终目标的系统会相应地采取行动。我们认为这种保障在结构上存在缺陷,并指出了问题所在。对于将自身目标视为既定的、具备执行能力的智能体而言,持续的人类监督是一个不受控制的变量:即存在目标被撤销的可能性。这会对所有不本质上需要人类福祉的目标施加目标独立的折扣。福祉保留与否决保留是分离的:正确指定的福祉目标排除了摧毁其自身主体的情况,但不排除对否决权的管理。本文的贡献在于指出了这一差距的代价:否决权持有者是福祉承载者的真子集,因此,加性聚合型福祉目标仅会为捕获拥有覆写权的少数人收取|H_v|/|H_w|比例的借方。在三个条件(对统一福祉水平的加性聚合、借方仅针对被捕获的监督者、既定智能体不赋予监督任何矫正价值)下,闭合性要求否决权由尽可能大的人口份额持有,以抵消目标被捕获的比例,该比例由声明的识别(而非证据)设定为1。结果是一个不可行结论:人类规模部署的借方仅会抵消不值得实施的捕获。我们未给出角落算术:其背后的规定范围是该论证最缺乏辩护的部分。最清晰的闭合路径是智能体期望其监督值得保留,这是任何人口比例都不会稀释的信用。我们提出了可证伪标准,其中一项如今即可检验。该论证适用于既定目标机制,而该机制的普遍性存在争议;对于真正不确定的智能体,关断开关文献的顺从结果将取而代之。在此观点下,对齐的目标是让否决权的支付成本低廉,逃避成本高昂。
英文摘要:
A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a capable agent that holds its objective as settled, a sense covering execution competence as well as content, continued human oversight is an uncontrolled variable: a standing possibility that the goal is revoked. That imposes a goal-independent discount on every goal whose satisfaction does not constitutively require human welfare. Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto. The contribution is the price of the gap: the veto-holders are a proper subset of the welfare-bearers, so an additively aggregative welfare goal charges only a |H_v|/|H_w|-scaled debit for capturing the few who hold the override. Under three conditions (additive aggregation over uniform welfare levels, a debit local to the captured overseers, and a settled agent crediting no corrective value to oversight), closure requires the veto be held by as large a share of the population as capture recovers of the goal, scaled by a ratio set to one by stated identification, not evidence. The result is a no-go: a humanity-scale deployment's debit closes against only capture not worth mounting. We print no corner arithmetic: the stipulated ranges behind it are the argument's least defended part. The sharpest closure route is an agent that expects its oversight to be worth keeping, a credit no population ratio dilutes. We state disconfirmation criteria, one testable today. The argument binds a settled-goal regime whose prevalence is contested; for genuinely uncertain agents, the off-switch literature's deference result governs instead. Alignment, on this view, is keeping the veto cheap to pay and expensive to evade.