发表机构
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出宪法引导的水印框架,通过自然语言原则和离线优化为每个请求选择合适的水印配置,在鲁棒性优先场景下将改写后检测率提升最多14个百分点,同时保持总体质量与干净检测性能。
AI 中文摘要
水印技术使语言模型提供商能够识别由其模型生成的文本。然而,其期望属性可能相互冲突(即更强的水印信号会降低文本质量),而抵抗编辑的设计也可能助长伪造。提供商通过选择平衡竞争目标或优先考虑特定属性的配置来解决这些权衡。这两种方法都对具有不同需求的请求施加了共享的操作点,可能在措辞保留至关重要的场景中牺牲质量,或在可靠归因至关重要的场景中牺牲鲁棒性。为了实现灵活且可适应的设计,我们引入了宪法引导的水印(Constitution-Guided Watermarking),这是一个框架,根据提供商的要求(以自然语言原则列出)为请求选择适当的权衡。离线阶段,一个预训练推理代理检查宪法规则和水印实现,并利用经验反馈迭代地细化特定规则的配置。在部署阶段,一个独立的监控器识别适用的规则并检索相应的策略,包括水印豁免,而无需修改服务模型。此外,我们的框架支持离线并行优化和基于不断演变的提供商要求细化特定规则的配置,而不影响部署,并将每个部署的配置与其评估证据绑定,使部署决策可审计。在使用KGW和五条规则宪法的概念验证评估中,我们的框架选择响应提供商优先级的配置,并在鲁棒性优先的请求上将改写后检测性能比固定配置提高最多14个百分点,同时在名义0.1%假阳性率下,在总体质量和干净检测方面达到或超过所有基线。
英文摘要
Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (\ie~stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations that balance competing objectives or prioritize particular properties. Either approach imposes a shared operating point on requests with different requirements, potentially sacrificing quality where wording preservation matters or robustness where reliable attribution is essential. To allow flexible and adaptable designs, we introduce \emph{Constitution-Guided Watermarking}, a framework that selects request-appropriate trade-offs from provider requirements, listed as natural-language principles. \emph{Offline}, a pretrained reasoning agent examines constitutional rules alongside watermark implementations and iteratively refines rule-specific configurations using empirical feedback. \emph{At deployment}, a separate monitor identifies applicable rules and retrieves the corresponding policy, including watermarking exemptions, without modifying the serving model. Furthermore, our framework supports offline parallel optimization and refinement of rule-specific configurations based on evolving provider requirements without affecting deployment, and binds each deployed configuration to its evaluation evidence, making deployment decisions auditable. In a proof-of-concept evaluation using KGW and a five-rule constitution, our framework selects configurations responsive to provider priorities and improves post-paraphrase detection on robustness-prioritized requests by up to $14$ percentage points over fixed configurations, while matching or exceeding all baselines in aggregate quality and clean detection at a nominal $0.1\%$ false-positive rate.
CommentsWorking paper (under review)