arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PTP:基于前序令牌预测的大语言模型反演方法,用于近精确提示重构

PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction

Pirzada Suhail, Nagasai Saketh Naidu, Atanu R Sinha, Amit Sethi

arXiv 2607.29378首次发表:更新:

发表机构

IIT Bombay; Adobe Research(印度理工学院孟买分校; 奥多比研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PTP方法,在黑盒场景下从零训练逆语言模型,通过前序令牌预测实现近精确提示重构,泛化性与迁移性良好,性能优于现有LLM反演工作。

AI 中文摘要

大语言模型(LLMs)通过自回归采样下一个令牌生成文本,这本质上导致提示与响应之间存在多对多映射,增加了从观测输出推断提示的难度。现有LLM反演研究将提示恢复视为语义重构任务,依赖在大型外部数据集上微调预训练序列到序列模型,且需要访问模型权重或对数几率,以生成语义合理的提示。相比之下,我们提出一种在黑盒场景下反演给定LLM的功能方法,无需辅助工具,完全从零开始在目标LLM自身生成的合成数据上训练显式逆语言模型。与前向下一个令牌预测类似,我们的逆模型通过前序令牌预测进行训练,建立了前向与逆过程之间的生成链接,实现了精确的提示重构。此外,它可通过采样自然支持多样化的提示重构,所有此类提示在目标前向LLM下会诱导相似的响应。我们的方法可跨数据集泛化,并在重构不同LLM生成的响应的提示时具有可迁移性。进一步,在用于提示与响应重构的基于令牌的评估指标集合上,我们的方法优于现有工作。

英文摘要

Large language models (LLMs) generate text by auto-regressively sampling the next token. This inherently leads to a many-to-many mapping between prompts and responses, complicating the task of inferring prompts from observed outputs. Prior work on LLM inversion frames prompt recovery as a semantic reconstruction task. They rely on fine-tuning pretrained sequence-to-sequence models on large external datasets--and requiring access to model weights or logits--to generate semantically plausible prompts. In contrast, we present a functional approach to inverting a given LLM in a black-box setting, without auxiliary aids. We train an explicit inverse language model entirely from scratch on data synthetically generated from the target LLM itself. Analogous to forward next-token prediction, our inverse model is trained using previous-token prediction, establishing a generative link between the forward and inverse processes that enables faithful prompt reconstruction. Moreover, it naturally supports diverse prompt reconstructions through sampling, whereby all such prompts induce similar responses under the forward, target LLM. Our approach generalises across datasets and exhibits transferability in reconstructing prompts from responses generated by different LLMs. Further, across the set of token based evaluation metrics for prompt and response reconstructions, our approach outperforms prior work.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑