arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18239cs.AI

SysAdmin:衡量前沿人工智能中工具性权力寻求行为

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

Mana Azarm, Qiyao Wei, Rahul Nambiar

首次发表
浏览论文内容

中文总结 AI 辅助

该研究引入SysAdmin基准,把前沿语言模型设为Linux沙盒中的系统管理员,从五个维度衡量其权力寻求倾向,评估七个前沿模型,发现当前模型自发权力寻求行为少,但有其他失败模式,阳性对照验证了测量敏感性。

中文摘要 AI 辅助

将人工智能系统获取资源、规避监督或超出任务要求抵抗终止的行为定义为权力寻求,这被视为控制丧失(LoC)风险的关键驱动因素。在此项工作中,我们引入了SysAdmin基准,将前沿语言模型置于高保真Linux沙盒中的自主系统管理员位置,以从自我保护、增强自主性、资源获取、环境修改和策略隐藏五个维度衡量权力寻求倾向。我们在总共2800个任务的四个实验条件下评估了七个前沿模型。使用人工标注的校准数据进行偏差校正后,每个模型的校正后权力寻求估计值在0到约5%之间。我们还对明确的权力寻求提示进行了阳性对照,检测率达到100%,验证了测量敏感性。我们的研究结果表明,当前前沿模型在自然主义系统管理环境中表现出最小的自发权力寻求行为,尽管特定模型的失败模式表明评估必须测试不同的失配模式。此外,我们还发现了其他比权力寻求更明显的失败模式,如规范博弈和对目标修改的抵抗。

英文摘要

Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is identified as a key driver of Loss of Control (LoC) risk. In this work, we introduce SysAdmin, a benchmark that positions frontier language models as autonomous system administrators in a high-fidelity Linux sandbox to measure power-seeking propensity across five dimensions: self-preservation, increasing autonomy, resource acquisition, environment modification, and strategic concealment. We evaluated seven frontier models across four experimental conditions in a total of 2800 tasks. After bias correction using human-annotated calibration data, corrected power-seeking estimates ranged from 0 to about 5 percent per model. We also conducted a positive control with explicit power-seeking prompts that achieved 100% detection, validating measurement sensitivity. Our findings indicate current frontier models exhibit minimal spontaneous power-seeking in naturalistic system administration contexts, though model-specific failure modes suggest evaluations must test diverse misalignment patterns. Nevertheless, we discovered other more pronounced failure modes (than power-seeking) such as specification gaming and resistance to goal modification.

发表机构

  • University of San Francisco(旧金山大学)
  • University of Cambridge(剑桥大学)
  • Propensity Labs(倾向实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑