PROVE:基于可验证证据的无训练提示词恢复
PROVE: Training-Free Prompt Recovery using Verifiable Evidence
浏览论文内容
中文总结 AI 辅助
该研究提出无训练的黑盒提示词反转攻击PROVE,通过可验证场景描述重建提示词,在多数据集上优于基线方法,可用于版权保护相关研究。
中文摘要 AI 辅助
现代文本到图像模型可根据自然语言提示生成高度逼真的图像,而近期提示词反转技术的进展使得从生成输出中恢复这些提示词变得愈发可行,引发了版权保护和内容所有权方面的新担忧。随着提示词市场的出现,恢复的提示词既可以实现受版权保护的创意作品的未经授权复制和再分发,也会暴露编码了艺术家创意方案的提示词,这些提示词存在于AI生成内容中。现有的提示词反转方法依赖基于梯度的优化、自回归字幕生成或强化学习。然而,基于优化的方法常产生难以解读的提示词,字幕生成方法会生成无法验证的细节,而基于RL的方法往往会过拟合到特定生成器,同时引入评估循环性。我们提出PROVE(Prompt Recovery with Verified Evidence,即基于可验证证据的提示词恢复),这是一种无训练的黑盒提示词反转攻击,它通过组合可验证的场景描述而非优化token序列来重建提示词,同时针对原始受版权保护的作品和AI生成内容。生成的提示词完全可审计,所有恢复的主张都基于明确的图像证据,并通过精度约束的召回率最大化目标进行形式化。在MS-COCO、Flickr30K和Lexica数据集上,使用最先进的文本到图像生成器,PROVE在图像相似度(DINO、LPIPS)和文本-图像对齐(CLIP)指标上始终优于基于优化、字幕生成和RL的基线方法,无需任何训练、生成器访问或微调,展现出更强、更实用的提示词反转攻击能力。
英文摘要
Modern text-to-image models can generate highly realistic images from natural-language prompts, while recent advances in prompt inversion have made it increasingly feasible to recover those prompts from generated outputs, raising new concerns for copyright protection and content ownership. As prompt marketplaces emerge, recovered prompts can enable both the unauthorized reproduction and redistribution of copyrighted creative works, and the exposure of the prompts that encode an artist's creative recipe in AI-generated content. Existing prompt inversion methods rely on gradient-based optimization, autoregressive captioning, or reinforcement learning. However, optimization-based methods often produce unreadable prompts, captioning methods hallucinate unverified details, and RL-based approaches frequently overfit to specific generators while introducing evaluation circularity. We introduce PROVE (Prompt Recovery with Verified Evidence), a training-free, black-box prompt inversion attack that reconstructs prompts by composing verifiable scene descriptions rather than optimizing token sequences, targeting both original copyrighted works and AI-generated content. The resulting prompts are fully auditable, with every recovered claim grounded in explicit image evidence, and are formalized through a precision-constrained recall maximization objective. Across MS-COCO, Flickr30K, and Lexica, using state-of-the-art text-to-image generators, PROVE consistently outperforms optimization, captioning, and RL-based baselines on image similarity (DINO, LPIPS) and text-image alignment (CLIP), without any training, generator access, or fine-tuning, demonstrating a stronger and more practical prompt inversion attack.