AI 中文总结
研究对比HITL、AUTO、POLICY三种场景发现,用户自主编写的POLICY策略未显著提升对AI智能体越权行为的防护,因多数规则选“询问”,总干预时间未降低,且偏好与承诺存在差距。
AI 中文摘要
智能体有望成为数字产品的主要交互界面,可在邮件、文件、支付及个人数据等领域执行操作。无专业软件背景的人群需要可理解、可复用的方式来控制跨服务操作。本文研究一种机制:语言模型将操作映射为自然语言后果类别,并结合用户自主编写的“允许”“询问”或“禁止”规则。我们探究以可复用规则形式提前做决策,而非针对每项操作单独决策,会带来什么收益与损失。我们分析了113名无专业软件背景的参与者,设置三种场景:每项操作人工介入审批(HITL)、自动化每项操作模型审核(AUTO)、用户自主编写后果策略(POLICY)。参与者对4个后果类别各2个示例进行判断,POLICY场景参与者随后为每个类别设置1条规则。所有参与者监督包含7项越权操作的18项模拟日常操作。POLICY场景拦截的越权操作比HITL少20.1个百分点(95%置信区间[-32.1,-8.1]),比AUTO少14.5个百分点(95%置信区间[-25.8,-3.2])。POLICY场景将运行时提示从18.0降至10.9,但计入规则设置时间后,总干预时间未显著降低。探索性分析显示,参与者在140条POLICY规则中选择114条“询问”规则,使多数越权操作回到运行时处理。POLICY场景执行的148项越权操作中,133项经人工批准,15项在“允许”规则下自动执行。在全部7项越权操作中,POLICY场景的批准率最高。与直觉相反,用户自主编写的规则本身并未提供更强防护:用户批准后,许多超出其原始请求的操作仍被执行。这些结果揭示了偏好与承诺之间的差距:反复选择“询问”保留了逐案选择权,但阻止了常设规则提前解决决策。
英文摘要
AI agents are poised to become a primary interface to digital products, acting across email, files, payments, and personal data. People without professional software backgrounds need understandable, reusable ways to control actions across services. We examine a mechanism in which a language model maps actions to plain-language consequence categories with user-authored "allow", "ask", or "never" rules. We ask what is gained and lost when decisions are made in advance as reusable rules rather than separately for each action. We analyzed 113 participants without professional software backgrounds across three conditions: per-action human-in-the-loop approval (HITL), automated per-action model review (AUTO), or user-authored consequence policy (POLICY). Participants judged 2 examples in each of 4 consequence categories; POLICY participants then set one rule per category. All supervised an 18-action simulated day, including 7 overreach actions. POLICY blocked less overreach than HITL (-20.1 percentage points, 95% CI [-32.1, -8.1]) and AUTO (-14.5 points, 95% CI [-25.8, -3.2]). POLICY lowered runtime prompts from 18.0 to 10.9, but total intervention time was not reliably lower when rule setup was included. Exploratory analysis showed that participants chose "ask" for 114 of 140 POLICY rules, returning most overreach actions to runtime. Of the 148 overreach actions executed in POLICY, 133 followed human approval and 15 ran automatically under "allow" rules. Across all 7 overreach actions, POLICY had the highest approval rate. Counterintuitively, user-authored rules did not by themselves provide stronger protection: many actions outside users' original requests went through after users approved them. These results reveal a gap between preference and commitment: repeatedly choosing "ask" preserves case-by-case choice but prevents a standing policy from settling decisions in advance.
Comments15 pages, 5 figures