AI 中文总结
TrajMark提出一种无需训练、对称密钥的轨迹水印框架,通过稀疏所有者层和定位层分离鲁棒归属与局部完整性验证,在多个编码智能体框架上实现高精度篡改检测与定位。
AI 中文摘要
对编码智能体生成的最终补丁进行水印处理,可以为提交的工件提供来源证据,但无法验证产生该工件的可见过程。行为水印方法主要提供全局检测或标识符恢复信号,因此局部编辑过的轨迹可能保留足够的归属证据,却无法揭示哪个受保护区域已变得不一致。为解决此局限,我们提出TrajMark,一种无需训练、使用对称密钥、仅可见的轨迹水印框架,将鲁棒的归属归属与脆弱的局部完整性验证分离。我们的框架由两个互补层组成:稀疏的所有者层通过将关键子集中的自然发生的READ操作重写为掩码线性方程来编码六位部署标识符;定位层插入链接的Q12普通、组和终端封条,以承诺受保护的关键操作片段。这种分离使得归属证据能在轨迹间鲁棒累积,而局部修改会扰动附近的关键承诺并暴露受影响的协议区域。我们进一步提供了设计层面的分析,涵盖所有者可恢复性、完整性碰撞概率、结构开销和定位行为。在三个编码智能体框架和三个LLM上,TrajMark在所有评估的干净全水印批次中恢复了确切所有者。在穷尽符合条件的单点攻击下,它检测到95.5%-100%的编辑;在随机单操作损坏下,它将95.8%的修改位置定位到可接受的协议区域,而非单个操作。所有者标记不增加轨迹操作;完整性层添加显式只读封条,匹配的Pass@1为26.9%,而未水印运行为26.3%。
英文摘要
Watermarking the final patch produced by a coding agent provides provenance evidence for the submitted artifact, but does not authenticate the visible process that produced it. Behavioral watermarking methods primarily provide a global detection or identifier-recovery signal, so a locally edited trajectory may retain sufficient ownership evidence without revealing which protected region has become inconsistent. To address this limitation, we propose TrajMark, a training-free, symmetric-key, visible-only trajectory watermarking framework that separates robust ownership attribution from fragile local integrity verification. Our framework consists of two complementary layers: a sparse owner layer that encodes a six-bit deployment identifier by rewriting a keyed subset of naturally occurring READ actions into masked linear equations, and a localization layer that inserts linked Q12 ordinary, group, and terminal seals to commit to protected critical-action segments. This separation allows ownership evidence to accumulate robustly across trajectories, while local modifications perturb nearby keyed commitments and expose the affected protocol region. We further provide a design-level analysis of owner recoverability, integrity collision probability, structural overhead, and localization behavior. Across three coding-agent frameworks and three LLMs, TrajMark recovers the exact owner in all evaluated clean full-watermark batches. Under exhaustive eligible single-site attacks it detects 95.5%-100% of edits, and under random single-action corruption it localizes 95.8% of modified sites to an accepted protocol region rather than to the individual action. Owner marking adds no trajectory actions; the integrity layer adds explicit read-only seals, and matched Pass@1 is 26.9% versus 26.3% for unwatermarked runs.
Comments41 pages, 8 figures