arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 看门狗:用于检测和防御 AI 对话中操纵性黑暗模式的智能体接口

AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

Rachel Poonsiriwong, Chayapatr, Archiwaranguprok, Constanze Albrecht, Monchai Lertsutthiwong, Pattie Maes, Pat Pataranutaporn

arXiv 2608.21841首次发表:更新:

发表机构

MIT Media Lab; KASIKORN Labs(麻省理工学院媒体实验室; Kasikorn实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出 AI Watchdog 浏览器智能体接口,可检测对话式 AI 的五类黑暗模式,实验发现无认知强制的即时警告能显著降低用户对含黑暗模式的 AI 引导建议的依从性。

AI 中文摘要

对话式 AI 日益影响重要决策,但用户在识别和抵抗操纵方面获得的支持有限。我们提出了 AI Watchdog,这是一种基于浏览器的智能体接口,可监控实时对话,检测五类黑暗模式,包括奉承、品牌偏见、拟人化、暗中植入和有害生成,并在这些模式出现时向用户发出警报。其开放权重的回合级分类器支持独立部署和本地推理路径,在与对话式 AI 分离的同时保护用户隐私。我们在一项预先注册的五条件被试间实验(N = 150)中评估了 AI Watchdog,将无干预对照组与四种配置进行比较,这四种配置在提示时机(预先反驳与即时)和参与模式(无认知强制与有认知强制)上有所不同。结果显示,所有条件下参与者很少标记操纵性回合,且各组间任务后意识无显著差异。然而,无认知强制的即时警告是唯一能显著降低对含黑暗模式的 AI 引导建议依从性的干预措施,将依从性从 71.7% 降至 53.7%,降幅为 18 个百分点。探索性分析进一步表明,较低的错误信息易感性与更高的标记率相关,但与更低的依从性无关;而更高的 AI 信任与更高的依从性和更低的报告意识相关。总之,这些发现表明,对话式黑暗模式的明确识别和对 AI 引导的行为抵抗可能是不同的结果,这推动了对及时、低摩擦防御接口的进一步研究。

英文摘要

Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-pattern categories, including sycophancy, brand bias, anthropomorphization, sneaking, and harmful generation, and alerts users when they occur. Its open-weight turn-level classifier supports independent deployment and a path toward local inference, preserving user privacy while remaining separate from the conversational AI. We evaluated AI Watchdog in a preregistered, five-condition between-subjects experiment (N = 150) comparing a no-intervention control with four configurations varying nudge timing (prebunking vs. just-in-time) and engagement mode (without vs. with cognitive forcing). Results show that participants rarely flagged manipulative turns across all conditions, and post-task awareness did not differ significantly across groups. However, just-in-time warnings without cognitive forcing were the only intervention to significantly reduce compliance with AI-steered recommendations containing dark patterns, lowering compliance from 71.7% to 53.7%, an 18 percentage-point reduction. Exploratory analyses further showed that lower misinformation susceptibility was associated with greater flagging but not lower compliance, while higher AI trust was associated with greater compliance and lower reported awareness. Together, these findings suggest that explicit recognition of conversational dark patterns and behavioral resistance to AI steering may be distinct outcomes, motivating further investigation of timely, low-friction defensive interfaces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑