发表机构
Chalmers University of Technology; University of Gothenburg; The University of Melbourne; Eindhoven University of Technology(查尔姆斯理工大学; 哥德堡大学; 墨尔本大学; 埃因霍温理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过访谈和调查软件从业者,探究其对智能体AI助手行为的理解及权限决策机制,提出权限系统应改进审查、区分意图与许可、明确可逆性等设计建议。
AI 中文摘要
智能体AI助手越来越多地代表开发者采取行动,例如修改文件、执行命令和访问外部资源。这些操作通常需要权限,但关于从业者如何在受益于智能体自主性的同时做出权限决策,目前知之甚少。为弥补这一空白,我们开展了一项顺序混合方法研究,首先访谈了18名使用AI智能体的从业者,然后基于访谈结果对115名从业者进行了问卷调查。我们发现,从业者通常通过他们可以直接观察和审查的内容来理解智能体行为,而决策、数据使用及其他幕后活动则不太清晰。这种不确定性也影响了权限决策,这些决策取决于操作的范围和风险、是否符合任务、对智能体的熟悉程度以及其运行环境。从业者通过调整对智能体的监督程度来应对,从预先设定限制到监控执行过程并在事后审查工作。他们施加的审查程度取决于信任、任务重要性、时间压力以及操作后果等因素。我们的研究结果表明,权限系统应使有重大后果的操作更易于审查,区分智能体被允许做什么与用户意图做什么,使可逆性更加清晰,避免将重复批准视为稳定偏好,并区分拒绝单个操作与拒绝整个方法。
英文摘要
Agentic AI assistants increasingly act on developers' behalf by modifying files, executing commands, and accessing external resources. These actions often require permission, yet little is known about how practitioners make permission decisions while still benefiting from agent autonomy. To address this gap, we conducted a sequential mixed methods study, interviewing 18 practitioners who use AI agents and then surveying 115 practitioners based on the interview findings. We find that practitioners often understand agent behaviour through what they can directly observe and review, while decisions, data use, and other activity behind the scenes remain less clear. This uncertainty also shapes permission decisions, which depend on the scope and risk of an action, whether it fits the task, familiarity with the agent, and the environment in which it operates. Practitioners respond by adjusting how closely they oversee agents, from setting limits in advance to monitoring execution and reviewing work afterwards. How much scrutiny they apply depends on factors such as trust, task importance, time pressure, and the consequences of an action. Our findings suggest that permission systems should make consequential actions easier to review, distinguish what an agent is allowed to do from what the user intended, make reversibility clearer, avoid treating repeated approvals as stable preferences, and distinguish rejecting a single action from rejecting an entire approach.
Comments21 pages, 6 figures