arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于硬提示(Hard Prompt)的能力与局限

On the Capability and Limitation of Hard Prompt

Lijia Yu, Shuaitong Liu, Gaojie Jin, Xinyu Li, Xiao-Shan Gao

arXiv 2609.32302首次发表:更新:

AI 中文总结

本文首次从理论上系统研究硬提示(离散提示)的能力与局限,证明其存在性判定为NP完全、最优求解为NP难,揭示其不完备性及长提示的“提示主导答案”现象,并给出泛化性的充要条件,为实际使用提供可靠指导。

AI 中文摘要

提示工程已成为使用大型语言模型(LLMs)不可或缺的工具,它能在不改变模型权重的情况下将LLMs转变为特定任务的专家。尽管提示工程在理论上已取得显著进展,但针对更实用的硬提示(即离散提示)的理论在很大程度上仍是空白。本文试图填补这一空白,要么提供完整的解决方案,要么在关于硬提示的三个核心理论问题上取得实质性进展。首先,我们证明:判断是否存在一个硬提示使Transformer解决下游任务的问题是NP完全的,而寻找最优硬提示是NP难的,据我们所知,这是关于硬提示的第一个计算复杂性结果。其次,我们表明,与软提示(即连续提示)不同,硬提示存在本质局限:硬提示不完备;短硬提示不能显著增强Transformer的能力;长硬提示表现出“提示主导答案现象”,即对于所有相同长度的查询,高概率下会给出相同的答案。另一方面,线性硬提示则不具有短提示或长提示的上述局限。第三,我们针对提示在有限任务上的性能推广到整个数据分布的情况,提供了任务规模(以提示长度表示)的紧界,从而给出了可推广性的充要条件。据我们所知,这是关于提示泛化性的第一个结果。我们的发现不仅为硬提示提供了首个理论洞见,也为现实世界中的LLM使用提供了可证明可靠的实际指导。

英文摘要

Prompt engineering has become an indispensable tool for using large language models (LLMs), turning LLMs into task-specific experts without changing their weights. Despite notable theoretical advances in prompt engineering, the theory for the more practical hard or discrete prompts is largely open. In this paper, we try to fill this gap either by providing a complete solution or by making substantial progress on the three core theoretical questions regarding hard prompts. First, we show that determining the existence of a hard prompt for a transformer to solve a downstream task is NP-complete and that finding an optimal hard prompt is NP-hard, which is the first computational complexity result for hard prompting, as far as we know. Second, we show that, unlike soft or continuous prompts, hard prompts have essential limitations: hard prompts are not complete; short hard prompts do not significantly enhance the ability of transformers; and long hard prompts exhibit the "prompt dominating answer phenomenon," meaning that, with high probability, the same answer is given for all queries of the same length. On the other hand, linear hard prompts do not have the limitations of short or long prompts. Third, we provide a tight bound on the size of the task in terms of the prompt length for the performance of prompts on the finite task to generalize to the entire data distribution, leading to a necessary and sufficient condition for generalizability. This is the first result on generalization for prompting, as far as we know. Our findings not only offer the first theoretical insights into hard prompts but also provide provably reliable practical guidance for real-world LLM usage.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑