arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越全局规范:无需重新训练即可在语言模型中实现毒性敏感性个性化

Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining

Rares A. C. Diaconescu, Iulia Slanina, Alina Florea, Andrei B. Trache, Miruna E. Coroi, Anne Arzberger, Jie Yang, Enrico Liscio

arXiv 2607.23175首次发表:更新:

发表机构

Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何在语言模型中实现毒性敏感性个性化,通过对免训练方法在三个推理阶段进行比较评估,与PRISM数据集目标对比,各方法降低对齐误差28 - 47%,揭示了多目标间权衡。

AI 中文摘要

减少毒性通常被视为一个全局对齐问题,但对有害语言的认知是主观且依赖上下文的。我们首次对在三个推理时间干预阶段将语言生成与用户特定毒性敏感性对齐的免训练方法进行了比较评估,包括预解码、解码中和解码后阶段。与从PRISM数据集得出的毒性敏感性目标相比,所有方法都将对齐误差降低了28 - 47%。结果揭示了对齐有效性、个性化和一般语言质量之间的基本权衡,表明毒性敏感性对齐是一个固有的多目标问题。

英文摘要

Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language generation to user-specific toxicity sensitivities across three inference-time intervention stages: pre-decoding (prompt conditioning and rewriting), in-decoding (token, logit, and representation steering), and post-decoding (candidate re-ranking). Evaluated against toxicity sensitivity targets derived from the PRISM dataset, all methods reduce alignment error by 28-47%. However, the results reveal a fundamental trade-off between alignment effectiveness, personalization, and general language quality, showing how toxicity sensitivity alignment is an inherently multi-objective problem.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑