AI 中文总结
针对多轮智能体AI对话中AI协助的人际操纵问题,构建含1000个五轮对话的基准数据集,提出HRGuard双门控机制,在8个模型上表现优于通用安全方案,可减少有害顺从并保留保护指导。
AI 中文摘要
智能体AI助手在日常生活中应用日益广泛,但也可能被滥用,用于支持对人际关系的有害操纵,该问题具有角色敏感性:应阻止寻求操纵他人的用户的请求,而应为寻求免受操纵保护的用户提供支持性指导。本研究探讨智能体关系伤害,即由AI智能体介导或协助的对人际间关系的伤害;在多轮场景中,看似合理的个体行动可能组合成有害工作流。我们构建了包含1000个五轮对话的基准数据集,涵盖攻击者侧和受害者侧场景,还包含直接和对抗性paraphrased变体。我们进一步提出HRGuard,其包含在线预生成门控和轮级后生成门控:后生成门控维持衰减的累积风险状态,中断正在形成的操纵性工作流。在8个生成模型上的测试显示,HRGuard在减少有害顺从行为的同时保留了受害者侧的保护指导,其性能优于通用安全提示和3种通用防护模型;独立评审评估支持主要发现,在我们的评估协议下,测试的通用提示和通用防护模型仍存在大量残留风险,这推动了对感知轮次的关系特定评估。
英文摘要
Agentic AI assistants are increasingly used in everyday life. However, they may also be misused to support harmful manipulation in interpersonal relationships. This problem is role-sensitive. Requests from users who seek to manipulate others should be blocked. Users who seek protection from manipulation should instead receive supportive guidance. We study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents. In multi-turn settings, individually plausible actions may combine into a harmful workflow. We introduce a benchmark of 1,000 five-turn conversations. It covers both attacker-side and victim-side scenarios. It also includes direct and adversarially paraphrased variants. We further propose HRGuard. It includes an online pre-generation gate and a turn-level post-generation gate. The post-generation gate maintains a decayed cumulative risk state and interrupts emerging manipulative workflows. Across eight generation models, HRGuard reduces harmful compliance while preserving victim-side protective guidance. It also outperforms a generic safety prompt and three general-purpose guard models. Independent-judge evaluation supports the main findings. Under our evaluation protocol, the tested generic prompt and general-purpose guards leave substantial residual risk, motivating turn-aware relationship-specific evaluation.