发表机构
The Hong Kong University of Science and Technology (Guangzhou); Bosum Institute of Management Science(香港科技大学(广州); 博硕管理科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型文档暴露溯源问题,提出基于源的语义水印SemTrace,通过构建文档特定二进制签名实现模型不可知的受保护副本暴露检测。
AI 中文摘要
大语言模型越来越多地被用于读取文档并生成下游文本,当文档所有者无法控制或检查执行生成的模型时,就会产生溯源问题。我们提出SemTrace,一种基于源的语义水印,用于检测生成的评论是否受已知受保护手稿副本的影响。SemTrace并非偏向令牌概率或施加表面形式模式,而是从手稿本身直接支持的事实命题构建特定于文档的二进制签名。受保护的PDF不可见地携带内容合约,该合约从每对事实中选择一个事实,并要求遵循指令的评论者在固定评论槽中表达这些事实,且不改变其独立评估。然后,一个冻结的自然语言推理模型对生成的语义证据进行显式擦除解码,并将恢复的位与分配给该副本的码字进行评分。该设计旨在实现模型不可知的分配副本暴露检测,同时使水印在语义上与源文档绑定。
英文摘要
Large language models are increasingly used to read documents and produce downstream text, creating a provenance problem when the document owner cannot control or inspect the model that performs the generation. We introduce SemTrace, a source-grounded semantic watermark for detecting whether a generated review was influenced by a known protected manuscript copy. Rather than biasing token probabilities or imposing surface-form patterns, SemTrace constructs a document-specific binary signature from factual propositions that are directly supported by the manuscript itself. A protected PDF invisibly carries a content contract that selects one fact from each binary pair and asks an instruction-following reviewer to express those facts in fixed review slots without changing its independent evaluation. A frozen natural language inference model then decodes the resulting semantic evidence with explicit erasures and scores the recovered bits against the codeword assigned to that copy. This design targets model-agnostic, assigned-copy exposure detection while keeping the watermark semantically tied to the source document.