arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CompanionHarm:用于检测现实世界AI伴侣对话中危害的多轮基准数据集

CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

Renwen Zhang, Han Meng, Jian Chai, Yuntao Lin, Yi-Chieh Lee

arXiv 2608.25377首次发表:更新:

AI 中文总结

本研究推出CompanionHarm基准数据集,含2111段用户与Replika的多轮对话,标注了13类危害,评估7个LLM发现多轮上下文检测危害更优,为相关研究提供基础。

AI 中文摘要

随着AI伴侣日益融入日常生活,迫切需要检测在社会情感层面的人机交互中出现的危害。然而该领域的研究因缺乏用于定义和评估具有关系性与情境性的危害的现实世界多轮对话数据集而受到限制。本研究中,我们推出CompanionHarm,这是一个公开可用的基准数据集,包含用户与AI伴侣Replika之间的2111段现实世界多轮对话(共14051轮发言)。7016轮AI发言由三名标注者基于AI伴侣危害分类体系,针对13类危害行为独立标注,数据集同时包含聚合标签与标注者层面的标签,以支持模型评估与系统性分歧分析。对七个大型语言模型(LLM)的评估显示,利用多轮对话上下文进行危害检测的效果优于基于孤立发言的检测,不过当前LLM仍难以一致地整合上下文线索、校准危害严重程度以及解读关系边界。我们还发现,对于依赖上下文的危害行为,标注者存在显著分歧,且分歧程度随标注者的政治派别、对话长度及发言位置而变化。总体而言,CompanionHarm为检测多轮人机对话中的社会情感危害,以及严谨探究人类与LLM如何解读此类危害提供了基础。我们的数据集可通过该https链接获取。

英文摘要

As AI companions become increasingly embedded in everyday life, there is an urgent need to detect harms that emerge in social and emotional human-AI interactions. Yet research in this area is constrained by the lack of real-world, multi-turn conversational datasets for operationalizing and evaluating harms that are relational and contextual. In this work, we introduce CompanionHarm, a publicly available benchmark dataset comprising 2,111 real-world, multi-turn conversations (14,051 utterances) between users and the AI companion Replika. 7,016 AI utterances were annotated independently by three annotators across 13 harmful behavior categories grounded in a taxonomy of AI companion harms, and the dataset includes both aggregated labels and annotator-level labels to support model evaluation and systematic disagreement analysis. Evaluations of seven large language models (LLMs) show that harm detection using multi-turn conversational context outperforms detection based on isolated utterances, although current LLMs still struggle to consistently integrate contextual cues, calibrate harm severity, and interpret relational boundaries. We also find substantial annotator disagreement for context-dependent harmful behaviors, with disagreement varying according to annotators' political affiliation, conversation length, and the utterance's position. Together, CompanionHarm provides a foundation for detecting socio-emotional harms in multi-turn human-AI conversations and for rigorously examining how such harms are interpreted by both humans and LLMs. Our dataset is available at https://github.com/HanMeng2004/CompanionHarm.

Comments10 pages, this dataset is publicly available at https://github.com/HanMeng2004/CompanionHarm

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑