SDARE-Bench:评估大型语言模型在二元与群体对话中的对话式污名检测与响应能力
SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue
浏览论文内容
中文总结 AI 辅助
研究人员构建首个场景式基准SDARE-Bench,评估8个LLMs在二元与群体对话中的污名检测和响应能力,发现LLMs对污名识别差、群体场景污名表达率达97.5%,污名响应是其安全漏洞。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地被用于可能影响社会判断的建议寻求和决策场景中。尽管污名对个人和群体有深远影响,但相关基准仍十分稀缺。现有通用领域评估通常依赖静态提示和固定格式任务,忽略了日常交流中的对话语境和受众效应。为解决这些差距,我们推出SDARE-Bench,这是首个基于场景的基准,用于评估大型语言模型的污名检测和开放式响应生成能力,包含1138个二元查询和1388个群体对话。对8个大型语言模型的实证结果一致显示,它们对污名成分的识别能力较差,尤其是在群体对话中。在开放式响应生成方面,群体场景下的污名表达显著高于二元场景,对污名的抵抗更弱,且建议更不切实际。响应使用在1392个人类标注响应上训练的分类器进行评估。在构建的群体压力场景中,污名表达率进一步升至惊人的平均97.5%。我们的研究结果表明,污名响应是大型语言模型反复出现的安全漏洞,尤其在社交复杂的对话语境中。
英文摘要
Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound effects on people and communities, benchmarks remain scarce. Existing general-domain evaluations typically rely on static prompts and fixed-format tasks, overlooking conversational contexts and audience effects in everyday communication. To address these gaps, we introduce SDARE-Bench, the first scenario-based benchmark evaluating both stigma detection and open-ended response generation in LLMs, comprising 1,138 dyadic queries and 1,388 group dialogue. Empirical results across 8 LLMs consistently demonstrate poor identification of stigma components, especially in group dialogues. In open-ended response generation, stigma expression was substantially higher in group settings than in dyadic, with weaker resistance to stigma and more unrealistic advice. Responses were evaluated using a classifier trained on 1,392 human annotated responses. In constructed group pressure settings, stigma expression rates further increased to a striking average of 97.5%. Our findings identify stigma response as a recurring LLM safety vulnerability, especially in socially complex conversational contexts.
发表机构
- Monash University(莫纳什大学)
- University of Liverpool(利物浦大学)
- University of Edinburgh(爱丁堡大学)
- Federation University(联邦大学)
- Orygen, The University of Melbourne(墨尔本大学下属Orygen机构)
机构由 AI 辅助整理,请以论文原文为准。