arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17713cs.CL

AEGIS:迭代保障的意识增强引导

AEGIS: Awareness-Enhanced Guidance for Iterative Safeguard

Kyungwon Park, Sangmin Lee, Heejae Chon, Hyungu Kang

首次发表
浏览论文内容

中文总结 AI 辅助

研究文本解毒中跨度引导的作用,提出AEGIS框架,结合跨度级检测器输出与生成器主干,分析其对毒性降低和意义保留平衡的影响,发现跨度引导解毒有条件有用,受生成器主干和语言上下文影响。

中文摘要 AI 辅助

在文本解毒中,跨度级别的基本原理通常被认为可提高可控性,但尚不清楚这种引导何时有帮助以及何时会带来权衡。我们提出了迭代保障的意识增强引导(AEGIS),作为研究跨英语、汉语普通话和韩语的跨度引导多语言解毒的探索性框架。AEGIS将跨度级检测器输出与冻结的生成器主干相结合,在重写过程中允许将有害跨度、强度标签和目标属性作为结构化引导提供。我们分析了跨度引导如何影响跨生成器家族、模型规模和语言的毒性降低与意义保留之间的平衡。结果表明跨度引导的解毒有条件地有用:明确的基本原理改变了毒性降低与意义保留之间的权衡,但其效果强烈依赖于生成器主干和语言上下文。这些发现凸显了跨度级控制信号在多语言解毒中的前景和局限性。

英文摘要

Span-level rationales are often assumed to improve controllability in text detoxification, but it remains unclear when such guidance helps and when it introduces trade-offs. We present Awareness-Enhanced Guidance for Iterative Safeguard (AEGIS) as an exploratory framework for studying span-guided multilingual detoxification across English, Mandarin Chinese, and Korean. AEGIS combines span-level detector outputs with frozen generator backbones, allowing harmful spans, intensity labels, and target attributes to be provided as structured guidance during rewriting. Rather than claiming state-of-the-art detoxification performance, we analyze how span guidance affects the balance between toxicity reduction and meaning preservation across generator families, model scales, and languages. Our results suggest that span-guided detoxification is conditionally useful: explicit rationales change the trade-off between toxicity reduction and meaning preservation, but their effects depend strongly on the generator backbone and the linguistic context. These findings highlight both the promise and the limitations of span-level control signals for multilingual detoxification.

发表机构

  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑