arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37310cs.CRcs.LG

AutoMark:实现自动研究以发现更优的LLM水印

AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks

  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Thibaud Gloaguen, Robin Staab, Martin Vechev

AI总结:

本工作提出AutoMark框架,首次实现自主发现LLM水印方案,通过严格标准和统计测试,利用前沿模型发现超50种方案,其中多种在可检测性、质量和鲁棒性上超越现有方法,并揭示新思想。

AI中文摘要:

随着LLM水印技术被商业部署并受到法规要求,提高其可靠性和有效性变得至关重要。然而,该领域近期的进展越来越多地依赖于改进现有方法的细节,这一努力从根本上受限于人类研究者的速度。在本工作中,我们首次实现了自主发现新的无失真、最先进的水印方案。为实现此目标,我们(i)建立了严格的标准以确保水印的可靠性(例如,它们不应具有意外高的假阳性率),(ii)提出了严格的统计测试来自动评估水印方案是否满足我们的标准,以及(iii)设计了一个评估套件,沿三个关键维度对水印进行排名:可检测性、质量和鲁棒性。通过使用3个前沿模型(GPT-6 Astra、Opus 5、Gemini-3.8 Flash)运行我们的框架,我们发现了超过50种不同的水印方案,其中几种在所有关键维度上均优于先前的工作。我们通过手动研究发现的方案进行补充,将关键思想提炼为较小的组件,并单独研究每个组件在各维度(可检测性、质量、鲁棒性)上的影响,以更好地理解所提方案的工作原理。重要的是,我们发现智能体在改进现有思想的同时,也发现了根本性的新思想(例如,将水印分数与随机每请求方向对齐)。总体而言,我们的工作建立了完全自主水印研究的第一步,使得发现更可靠和有效的水印成为可能。我们的代码可在该https URL获取,博客文章可在该https URL可视化我们的结果。

英文摘要:

With LLM watermarking being deployed commercially and now required by regulations, improving its reliability and effectiveness has become crucial. Yet, recent progress in the field of LLM watermarking has increasingly been driven by improving details of existing methods, an effort fundamentally limited by the pace of human researchers. In this work, we enable for the first time the autonomous discovery of new distortion-free state-of-the-art watermarking schemes. To enable this, we (i) establish strict criteria to ensure that watermarks are reliable (e.g., they do not have an unexpectedly high false positive rate), (ii) propose rigorous statistical tests to automatically evaluate whether a watermarking scheme satisfies our criteria, and (iii) design an evaluation suite to rank watermarks along three key dimensions: detectability, quality, and robustness. By running our framework with 3 frontier models (GPT-6 Astra, Opus 5, Gemini-3.8 Flash), we discover over 50 different watermarking schemes, including several that outperform prior works along all key dimensions. We complement this by a manual study of the discovered schemes, distilling the key ideas into smaller components, and individually studying the impact of each component across dimensions (detectability, quality, robustness) to better understand how the proposed schemes operate. Importantly, we find that the agents, on top of improving existing ideas, also discover fundamentally new ideas (e.g., aligning watermark scores with random per-request direction). Overall, our work establishes the first steps of fully autonomous watermarking research, enabling the discovery of more reliable and effective watermarks. Our code is available at https://github.com/eth-sri/automark, and a blogpost to visualize our results at https://www.sri.inf.ethz.ch/blog/automark.

补充信息

↑