arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应协同服务:现代推理引擎上的大语言模型水印技术

Adaptive Co-Serving LLM Watermarking on Modern Inference Engines

Kieu Dang, Phung Lai, Ching-Yun Ko, Pin-Yu Chen

arXiv 2610.03955首次发表:更新:

发表机构

State University of New York at Albany; IBM(纽约州立大学奥尔巴尼分校; IBM)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SWIFT框架,将LLM水印与推理引擎协同设计,集成生成与水印构建,通过异步协同服务和自适应调度降低延迟,提升实用性、可检测性和鲁棒性,显著优于现有基线。

AI 中文摘要

大语言模型(LLM)水印技术对于所有权验证和知识产权保护至关重要。然而,现有方法侧重于算法设计,而将LLM推理引擎视为独立组件。这种分离常常引入辅助模型或外部工具,增加了延迟和内存开销,同时限制了现代推理优化的使用。因此,部署差距依然存在:实用的水印技术必须在最小化服务开销的同时,保持实用性、可检测性和鲁棒性。为弥合这一差距,我们提出了SWIFT,一个将LLM水印技术与现代推理基础设施协同设计的框架。SWIFT(i)在共享的LLM后端内集成文本生成和水印构建,以减少重新计算、提高缓存复用并简化系统复杂性;(ii)使用带语义保护的指令引导候选生成,以保持事实一致性和语义;(iii)通过异步协同服务和自适应调度,使水印与生成并发执行,以减少延迟并提高资源利用率;(iv)利用vLLM优化来提升服务效率。跨领域和任务的广泛实验表明,SWIFT在实现强实用性、可检测性、鲁棒性和下游性能的同时,显著提高了服务效率。具体而言,它实现了最高的实用性得分(4.87),检测准确率达99.65%,而低延迟基线SynthID为76.65%,在水印移除攻击下检测率高达97.7%,且延迟比最高实用性基线SafeSeal低5.9倍。这些结果证明了将水印技术与推理基础设施协同设计对实用LLM服务的益处。代码和工件可在以下网址获取:此https URL。

英文摘要

Large language model (LLM) watermarking is important for ownership verification and intellectual property protection. However, existing approaches focus on algorithmic design while treating LLM inference engines as separate components. This separation often introduces auxiliary models or external tools, increasing latency and memory overhead while limiting the use of modern inference optimizations. As a result, a deployment gap remains: practical watermarking must preserve utility, detectability, and robustness while minimizing serving overhead. To bridge this gap, we propose SWIFT, a framework that co-designs LLM watermarking with modern inference infrastructure. SWIFT (i) integrates text generation and watermark construction within a shared LLM backend to reduce re-computation, improve cache reuse, and simplify system complexity; (ii) uses instruction-guided candidate generation with semantic protection to preserve factual consistency and meaning; (iii) executes watermarking concurrently with generation through asynchronous co-serving and adaptive scheduling to reduce latency and improve resource utilization; and (iv) leverages vLLM optimizations to enhance serving efficiency. Extensive experiments across domains and tasks show that SWIFT achieves strong utility, detectability, robustness, and downstream performance while substantially improving serving efficiency. Specifically, it achieves the highest utility score (4.87), 99.65% detection accuracy compared with 76.65% for the low-latency baseline SynthID, and up to 97.7% detection under watermark removal attacks, with 5.9x lower latency than the highest-utility baseline, SafeSeal. These results demonstrate the benefits of co-designing watermarking with inference infrastructure for practical LLM serving. Code and artifacts are available at: https://anonymous.4open.science/r/SWIFT-0597.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑