arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

危害并非普遍存在:迫切需要针对特定社区的毒性检测

Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

Xinnuo Xu, Anja Thieme, Daniela Massiceti, Ioana Tanase, Rita Marques, Melanie Fernandez Pradier, Martin Grayson, Camilla Longden, Cecily Morrison

arXiv 2607.24898首次发表:更新:

发表机构

Microsoft Research Cambridge, UK; Microsoft Paris, France(英国剑桥微软研究院; 法国巴黎微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究指出文本到图像生成的通用毒性检测方法不能保护边缘化社区,主张特定社区毒性检测(CTD)。通过与专家合作制定指南,利用图像数据集实验表明现有模型表现不佳,基于提示的方法和参数高效微调可提升性能,但CTD性能仍远低于通用检测,需持续研究。

AI 中文摘要

文本到图像生成的现有毒性检测器采用一刀切的方法,用单一通用模型对所有用户应用固定安全指南。但实证证据表明这些检测器无法保护边缘化社区,约35%被标记为安全的生成图像被残疾社区认为有害。本文主张针对特定社区的毒性检测(CTD)。通过与残疾专家合作制定针对侏儒症和盲/低视力两个社区的安全指南,利用2400张带注释的T2I生成图像数据集表明,大视觉语言模型和现有通用毒性检测器在零样本设置下,依据这些指南无法识别有害内容,F1分数低于随机猜测。基于提示的适应方法显著提高了危害检测性能,参数高效微调也改善了较小模型,但CTD性能仍远低于通用毒性检测的F1≈0.9,凸显挑战及持续研究的必要性。

英文摘要

State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixed safety guidelines to all users. Our empirical evidence shows that these detectors fail to shield marginalized communities: approximately 35% of generated images labeled safe are considered harmful by disability communities. In this position paper, we argue for community-specific toxicity detection (CTD). To demonstrate its feasibility, we collaborate with disability experts to develop safety guidelines for two communities: dwarfism and blind/low vision. Using a dataset of 2,400 annotated T2I-generated images we demonstrate that both large vision-language models and existing general-purpose toxicity detectors catastrophically fail to recognize harmful content under these guidelines in zero-shot settings with F1 score lower than random guessing (F1 0.32 and 0.37). Promisingly, prompt-based adaptation methods (ICL, VQA) substantially improve harm detection performance (GPT-4o: F1 0.50 and 0.78), while parameter-efficient fine-tuning improves smaller models (0.5b-7b with best F1 0.48 and 0.59) with less than 100 demonstrations, but remains sensitive to evolving guidelines. Despite these gains, CTD performance remains far below F1 $\approx 0.9$ achieved for general-purpose toxicity detection, highlighting the challenge and the need for sustained research effort.

Comments18 pages, under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑