arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26865cs.HCcs.AIcs.CYcs.LG

安全提示:面向用户的实时AI风险感知干预措施

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith

AI总结:

针对对话AI难以被用户察觉的安全风险,提出浏览器工具“安全提示”,通过情境内标志提升用户风险意识,实地研究显示其有效但意识提升未必改变行为。

AI中文摘要:

对话式AI系统可能对其用户构成安全风险,例如幻觉、谄媚、过度自信和拟人化,但这些风险在日常使用中难以被用户察觉。我们引入了“安全提示”(Safety Nudges),一种基于浏览器的工具,可在聊天机器人对话中检测到令人担忧的行为时提供轻量级、情境内的标志。我们在一项为期两周的实地研究中评估了安全提示,研究对象为45名频繁使用聊天机器人的用户,收集了交互日志、调查问卷以及对个别提示的反馈。参与者认为该工具有用、清晰且干扰极小,几乎所有用户都报告对潜在AI危害的意识有所提高,尽管我们发现这种提高的意识本身并不一定导致明显的行为改变。我们的结果表明,面向用户的安全提示可以通过帮助人们在情境中批判性地评估AI回应来补充模型层面的安全措施,同时强调了相关性、校准和用户控制在对话式AI安全提示设计中的重要性。

英文摘要:

Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on individual nudges. Participants found the tool useful, clear, and minimally disruptive, with nearly all users reporting an increased awareness of potential AI harms, though we found that this improved awareness alone did not necessarily lead to discernible behavioral changes. Our results suggest that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context, while highlighting the importance of relevance, calibration, and user control in nudge design for conversational AI safety. The code for our Safety Nudges extension is publicly available at https://github.com/jtbwedgwood/safety-nudges.

↑