arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

助手还是执行者?使用通用人工智能代理时学生的信任、控制与委托遗憾

Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent

Shiva Pochampally, Shengwei An, Yan Chen

arXiv 2607.18257首次发表:更新:

发表机构

Department of Computer Science; Virginia Tech(计算机科学系; 弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨用户使用通用人工智能代理时的委托遗憾问题,通过大学生完成不同任务的对照研究,发现参与者按任务校准信任,不可逆与外部可见性影响信任撤回,代理无预览执行行动时委托遗憾常现,为代理设计提供了相关启示。

AI 中文摘要

当人工智能代理从回答问题转向采取行动时,用户面临新问题:决定委托何事给一个其行动空间无法完全预见的系统。我们将由此产生的不满称为委托遗憾,即用户后悔的并非代理犯错,而是其行为超出授权范围。在一项对照研究中,20名大学生使用通用人工智能代理OpenClaw完成五项日常任务,所选任务在隐私、风险和可逆性方面各异。针对每项任务,我们用5分量表测量信任、感知控制、透明度、监督负担和批准偏好,并收集通过主题编码分析的自由文本反馈。有三项发现:一是参与者按任务而非按代理校准信任;二是不可逆性与外部可见性共同作用而非仅风险因素导致信任撤回;三是代理无预览执行行动时委托遗憾始终存在,即便输出被评为成功。我们讨论了对代理设计的启示,包括明确行动边界、支持按任务的自主政策以及区分咨询输出与代理执行。

英文摘要

When AI agents shift from answering questions to taking actions, users face a new problem: deciding what to delegate, to a system whose action space they cannot fully anticipate. We call the resulting dissatisfaction delegation regret, a pattern in which users regret not that the agent erred, but that it acted beyond what they would have authorized. In a controlled study, 20 university students completed five common daily tasks using OpenClaw, a general-purpose AI agent, across tasks chosen to vary in privacy, stakes, and reversibility. For each task we measured trust, perceived control, transparency, supervision burden, and approval preference on 5-point Likert scales, and collected free-text reflections analyzed through thematic coding. Three findings emerged. First, participants calibrated trust per task rather than per agent: they granted wide autonomy for advisory and low-stakes tasks but demanded confirmation for irreversible, externally visible actions. Second, irreversibility combined with external visibility, rather than stakes alone, appeared to drive trust withdrawal: the moderate-stakes email task triggered the sharpest drop in trust (M = 3.10) and the highest demand for approval (M = 4.65), whereas a high-stakes but verifiable task did not produce the same response. Third, delegation regret appeared consistently when the agent executed actions without preview, even when the output was rated as successful. We discuss implications for agent designs that expose action boundaries, support per-task autonomy policies, and separate advisory output from agentic execution.

CommentsPresented at the 2026 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC). 10 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑