arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WaterMoE:基于专家路由的高保真高效水印技术

WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency

Z Sun, Q Jiang, S Sheng, L Xiang

arXiv 2607.13099首次发表:更新:

发表机构

Shanghai Jiao Tong University; Shanghai Innovation Institute; Tianjin University(上海交通大学; 上海创新研究院; 天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大型语言模型水印技术实践中模型性能下降和推理开销大的问题,提出WaterMoE方案,通过在专家选择中嵌入水印信号,在推理循环中实现水印嵌入,实验证明其保真性能好、优于现有方法且开销小,可用于实际任务。

AI 中文摘要

大型语言模型取得显著成功,但引发对内容来源和滥用的担忧,催生可靠水印技术需求。现有技术因模型性能严重下降和额外推理开销很少在实践中采用。为此构建基准评估9种代表性水印方法,发现现有方法多针对文本流畅性而非受限复杂任务,开销大。提出针对专家混合语言模型的水印方案WaterMoE,通过控制扰动在路由器的专家选择中嵌入水印信号,在推理循环中嵌入水印,质量下降和计算开销可忽略不计。实验表明该方法保真性能接近无水印模型,在基准测试中优于现有方法,加速4倍,推理延迟仅增加1%,证明可用于实际任务。

英文摘要

Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating the need for reliable watermarking techniques. However, these techniques have rarely been adopted in practice mainly for two reasons: i) severely degraded model performance, and ii) additional inference overhead. To confirm the problem, we construct a comprehensive benchmark spanning different generation tasks to systematically evaluate 9 representative watermarking methods. We found almost all existing methods are designed for text fluency, but not for restricted and complicated tasks, and their overhead prevents them from deployment in latency-critical systems. To address i) and ii), we propose an LLM watermarking scheme \textit{WaterMoE} for the growingly popular Mixture-of-Experts (MoE) LLMs. WaterMoE embeds watermarking signals through controlled perturbation into the expert selection at each router, which accumulates to token selection shift at the final output. In contrast to watermarking as a post-processing token-sampling approach, WaterMoE embeds watermark within the inference loop incurring negligible quality degradation and computational overhead. Extensive experiments demonstrate that our method achieves a fidelity performance close to the unwatermarked and consistently outperforms state-of-the-art watermarking methods on the benchmark, with up to $4\times$ speedup, incurring merely 1\% additional inference latency compared to native generation. The results demonstrate the capability of WaterMoE to be deployed in real-world tasks.

Comments21 page

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑