SEW:LLM生成代码的风格编码水印
SEW: Style-Encoded Watermarking of LLM-Generated Code
浏览论文内容
中文总结 AI 辅助
SEW通过风格规则、偏好校准和上下文聚合,在已生成的代码中嵌入和检测水印,在CodeContests上平均相对提升12.44%的TPR@FPR5%,并对代码编辑攻击保持鲁棒。
中文摘要 AI 辅助
代码水印支持对LLM生成的代码进行溯源追踪。在LLM生成代码时修改令牌选择以嵌入水印,会在可检测性与功能正确性之间产生权衡。其他方法则使用预定义的变换或经过训练的神经模型对已完成的代码进行水印处理。重复出现的模式可能使水印选择在程序间变得可预测,而将未加水印代码中常见的模式视为水印证据则可能导致误检。因此,我们引入了SEW,它通过三个组件在已生成的代码中嵌入和检测水印:(i)从风格指南和变换规则中收集的代码风格规则,风格选择由密钥和每个程序的结构上下文决定;(ii)风格偏好校准,使用从人类编写的代码中估计的风格概率来评估水印证据;(iii)上下文感知的风格聚合,结合来自结构匹配位置且被分配相同代码风格选择的证据,防止重复应用该选择导致水印证据膨胀。在CodeContests上,跨三个LLM和三种编程语言,SEW在TPR@FPR5%上相对于基线实现了12.44%的平均相对提升,并且对四种非LLM代码编辑攻击具有鲁棒性,平均相对下降仅为0.94%。我们的代码可在该https URL获取。
英文摘要
Code watermarking supports provenance tracking for code generated by LLMs. Modifying token selection to embed watermarks as an LLM generates code can create a trade-off between detectability and functional correctness. Other methods instead watermark completed code using predefined transformations or trained neural models. Recurring patterns can make watermark choices predictable across programs, while treating patterns common in unwatermarked code as watermark evidence can cause false detections. We therefore introduce SEW, which embeds and detects watermarks in already generated code through three components: (i) code style rules collected from style guides and transformation rules, with style choices determined by a secret key and each program's structural context; (ii) style-preference calibration, which evaluates watermark evidence using style probabilities estimated from human-written code; and (iii) context-aware style aggregation, which combines evidence from structurally matching locations assigned the same code style choice, preventing repeated applications of that choice from inflating watermark evidence. On CodeContests across three LLMs and three programming languages, SEW achieves a mean relative improvement of 12.44% in TPR@FPR5% over the baselines and is robust to four non-LLM code-editing attacks, with only a 0.94% mean relative decrease. Our code is available at https://github.com/suhanmen/SEW.
发表机构
- Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。