发表机构
University of North Carolina at Charlotte; George Mason University; Indiana University(北卡罗来纳大学夏洛特分校; 乔治梅森大学; 印第安纳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对模型归属分类器,提出ForgePrint框架,通过搜索-蒸馏方式改写文本以定向迁移模型指纹,在摘要任务上实现高目标成功率,揭示指纹检测不等于来源真实性。
AI 中文摘要
模型归属分类器通常能够识别出文本是由哪个语言模型生成的,这使得模型特有的写作模式成为来源的信号。然而,对未修改文本的准确归属并不能表明,在经过故意改写后,预测是否仍能识别出原始来源。我们将此问题表述为定向指纹迁移:改写一个模型的输出,使得归属分类器将其归因于一个选定的目标模型。我们研究摘要生成任务,其中不同模型接收相同的文档并表达相同的基础内容,为条件生成提供了一个受控环境。我们引入了ForgePrint,一个先搜索后蒸馏的框架,该框架首先搜索能够将归属向目标指纹移动的改写,然后将选定的改写蒸馏到一个单遍4B学生模型中。在CNN/DM上,学生模型达到了70.2%的目标成功率,优于其教师模型(54.1%)和六个已发布的改写基线中最强的(39.3%),且这些评估针对的是攻击从未查询过的保留分类器。在将摘要从一个开放模型转移到选定的商业模型时,它也达到了68.3%的目标成功率。这些结果表明,指纹可检测性不应与来源真实性混为一谈,并且仅基于文本的归属在定向改写下可能提供关于模型身份的误导性证据,即使它在未修改文本上是准确的。
英文摘要
Model-attribution classifiers can often identify which language model produced a text, making model-specific writing patterns a signal of provenance. Accurate attribution on unmodified text, however, does not show whether the prediction still identifies the original source after deliberate rewriting. We formulate this problem as targeted fingerprint transfer: rewriting one model's output so that attribution classifiers assign it to a chosen target model. We study summarization, where different models receive the same document and express the same underlying content, providing a controlled setting for conditional generation. We introduce ForgePrint, a search-then-distil framework that first searches for rewrites that move attribution toward a target fingerprint, then distils the selected rewrites into a one-pass 4B Student model. On CNN/DM, the Student reaches 70.2% target success rate, outperforming both its Teacher (54.1%) and the strongest of six published rewriting baselines (39.3%), against held-out classifiers that are never queried by the attack. It also reaches 68.3% target success when transferring summaries from an open model toward chosen commercial models. These results show that fingerprint detectability should not be conflated with source authenticity, and that text-only attribution can provide misleading evidence of model identity under targeted rewriting, even when it is accurate on unmodified text.
Comments31 pages, 7 figures, 25 tables. Project page: https://haohanyuan01.github.io/ForgePrint/