arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00361cs.CYcs.CL

净化有毒交流:负责任人工智能的设计科学方法

Detoxifying Toxic Communication: A Design Science Approach to Responsible AI

  • Schulich School of Business, York University(约克大学舒立克商学院)

机构由 AI 辅助整理,请以论文原文为准。

Hossein Arshadi Soufiani, Henry M. Kim, Hjalmar Turesson, Syed Mohammad Arham Noman, Anav Setia

AI总结:

本研究采用设计科学方法,构建含微调Transformer分类器与生成式净化模型的负责任AI,可检测并净化数字工作场所的有毒交流,相关设计原则兼顾意义保留与公平性。

AI中文摘要:

数字工作场所中的有毒语言,如侮辱性用语、讽刺、居高临下的态度及微妙的不文明行为,会侵蚀信任、士气与协作。现有审核工具主要通过删除或屏蔽有害消息来处理,这会中断交流且无法提供建设性解决方案。本研究采用设计科学研究方法,构建负责任人工智能制品,用于检测并净化有毒交流。该制品整合了微调后的基于Transformer的分类器(DistilBERT、DistilRoBERTa)与生成式净化模型(mT0-XL-Detox-ORPO),可将有毒文本重写为语义等价的非冒犯性释义。技术评估显示,其在毒性检测中准确率高,重写消息的语义保留性强,能在支持对话连续性的同时强化尊重性话语。本文还提出了负责任人工智能审核的设计原则,优先考虑意义保留与公平性。

英文摘要:

Toxic language in digital workplaces such as pejoratives, sarcasm, condescension, and subtle incivility can erode trust, morale, and collaboration. Existing moderation tools primarily delete or block harmful messages, disrupting communication and offering no constructive resolution. This study adopts a Design Science Research approach to create a responsible AI artifact that detects and detoxifies toxic communication. The artifact integrates fine-tuned transformer-based classifiers (DistilBERT, DistilRoBERTa) with a generative detoxification model (mT0-XL-Detox-ORPO) that rewrites toxic text into semantically equivalent, non-offensive paraphrases. Technical evaluation demonstrates high accuracy in toxicity detection and strong semantic preservation in rewritten messages, supporting conversation continuity while reinforcing respectful discourse. The paper contributes design principles for responsible AI moderation that prioritize meaning preservation and fairness.

补充信息

↑