Watermarking Degrades Alignment in Language Models: Analysis and Mitigation
水印会损害语言模型的对齐:分析与缓解
机构 * New Jersey Institute of Technology(新泽西理工学院)
专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.LG
AI总结 研究发现水印影响语言模型对齐,提出Alignment Resampling方法恢复对齐性能。
Comments Published in Transactions of Machine Learning Research 02/2026. Extended version of the earlier paper published at the 1st Workshop on GenAI Watermarking (ICLR 2025)
Journal ref Transactions of Machine Learning Research 02/2026