发表机构
Potsdam Institute for Climate Impact Research(波茨坦气候影响研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文设计了一种保留人类赋能的AI公平目标函数,通过公理化方法兼顾人类权力的不平等与风险规避,明确AI需赋能人类并管理人机权力平衡,还验证了其在典型场景的行为后果。
AI 中文摘要
本文探讨通过强制AI智能体明确赋能人类、以理想方式管理人与AI智能体间权力平衡,促进人机交互中人类福祉与安全的理念。采用基于理想性质的原则性部分公理化方法,设计了一种可参数化、可分解的AI系统目标函数,该函数代表对人类权力的不平等规避与风险规避型长期聚合。它可考虑人类有限理性与社会规范的模型,且关键在于涵盖人类的各类可能目标。我们证明部分理想性质如何强制特定函数形式并限制参数范围,在若干典型情境中举例说明软最大化该指标的后果,并描述其可能隐含的工具性子目标。
英文摘要
This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach based on desirable properties, we design a parametrizable and decomposable objective function for AI systems that represents an inequality- and risk-averse long-term aggregate of human power. It can take into account models of human bounded rationality and social norms, and crucially, considers a wide variety of possible human goals. We prove how certain desiderata enforce particular functional forms and restrict parameter ranges. We exemplify the consequences of softly maximizing this metric in several paradigmatic situations and describe what instrumental sub-goals it will likely imply.
CommentsSlightly extended version of paper accepted for Algorithmic Decision Theory 2026