基于跨度引导的解毒方法何时有效?受控比较中的人类偏好与评估者诊断
When Does Span-Guided Detoxification Help? Human Preferences and Evaluator Diagnostics in a Controlled Comparison
- Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究对比跨度引导与无引导解毒方法的人类偏好,发现二者存在权衡,不同分层偏好不同,自动评估无法替代人类判断,推动了新评估协议的发展。
AI中文摘要:
跨度引导的改写旨在通过将编辑操作限定在标注的有害跨度内来保留原意,但这种限制可能导致有害意图无法得到充分缓解。我们在包含人工精选输入和HateXplain测试项的混合源英语评估集上,对跨度引导与无引导解毒方法开展受控探索性比较。在固定单生成器设置下,我们开展密集盲法人类评估。人类偏好显示二者存在权衡,而非存在统一更优的改写策略:当局部编辑保留原立场且避免不必要修改时,跨度引导输出更受青睐;而当更广泛的改写实现更充分的缓解时,无引导输出更受青睐。这种对比在研究定义的分层中差异显著:在强分层中两种策略具有竞争力,而在温和分层中无引导改写明显更受偏好。原理标注将该差异归因于互补的失败风险:局部编辑后的残留危害,以及更广泛改写后的过度修改。我们将自动评估视为诊断工具而非人类判断的替代方案。毒性-相似度标量、多生成器分析及两种通用LLM评估器重现了部分整体趋势,但未产生类似的分层对比。这些特定设置的发现未建立基于严重程度的路由规则,反而推动了评估协议的发展,该协议需分别评估缓解充分性与原意保留情况,并报告残留危害、过度修改及整体得分。
英文摘要:
Span-guided rewriting aims to preserve meaning by localizing edits to annotated harmful spans, but the same constraint can leave harmful intent insufficiently mitigated. We present a controlled exploratory comparison of span-guided and unguided detoxification on a mixed-source English evaluation set comprising manually curated inputs and HateXplain test items. We conduct a dense blinded human evaluation under a fixed single-generator setting. Human preferences reveal a trade-off rather than a uniformly superior rewriting strategy. Span-guided outputs are favored when localized editing preserves the original stance and avoids unnecessary modification, whereas unguided outputs are favored when broader rewriting achieves more complete mitigation. This contrast varies substantially across the study-defined strata: the two strategies are competitive in the strong stratum, while unguided rewriting is clearly preferred in the mild stratum. Rationale annotations trace this difference to complementary failure risks: residual harm after localized editing and over-modification after broader rewriting. We treat automatic evaluation as a diagnostic rather than a substitute for human judgment. Toxicity-similarity scalarizations, a multi-generator analysis, and two general-purpose LLM judges reproduce parts of the aggregate tendency but do not yield an analogous stratified contrast. These setting-specific findings do not establish a severity-based routing rule. Instead, they motivate evaluation protocols that assess mitigation sufficiency and meaning preservation separately and report both residual harm and over-modification alongside aggregate scores.