arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11634cs.CRcs.AI

LTBD:用于提示注入防御的可学习信任边界分隔符

LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense

Luman Zhao, Minghui Xu, Yue Zhang, Yijun Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLMs的提示注入攻击问题,提出轻量级防御方法LTBD,通过可学习分隔符编码信任边界,在保持模型参数不变的情况下,防御效果优于推理时防御,接近训练式防御,且对自适应攻击有效。

中文摘要 AI 辅助

大型语言模型(LLMs)在复杂任务上表现出色,但仍极易受到提示注入攻击,即嵌入在外部数据中的恶意指令可覆盖用户意图。现有防御方法受限于模型微调需求、对自适应攻击的脆弱性,或依赖脆弱的手工提示。我们认为,这种脆弱性的根本来源是缺乏对信任来源的显式表示。为解决该问题,我们引入Learnable Trust-Boundary Delimiters(LTBD,可学习信任边界分隔符),这是一种轻量级防御方法,在保持LLM参数不变的同时,显式对输入中的信任边界进行编码。LTBD使用少量可学习分隔符区分可信用户指令与不可信外部数据,使模型能更好地遵循预期的信任层级。实验结果显示,LTBD的性能显著优于推理时防御方法,且与基于训练的方法表现相当,同时保留了良性任务的效用并引入可忽略的推理开销。具体而言,LTBD在AlpacaFarm上实现0.00%的攻击成功率(ASR),在TaskTracker上仅实现0.11-0.19%的ASR。LTBD在自适应攻击下仍保持有效,在此类攻击中,攻击者完全了解防御机制并明确试图绕过它。

英文摘要

Large language models (LLMs) perform remarkably well on complex tasks, yet remain highly vulnerable to prompt injection attacks, where malicious instructions embedded in external data can override user intent. Existing defenses remain limited by model fine-tuning requirements, vulnerability to adaptive attacks, or reliance on brittle handcrafted prompts. We argue that a fundamental source of this vulnerability is the lack of an explicit representation of trust provenance. To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged. LTBD uses a small number of learnable delimiters to distinguish trusted user instructions from untrusted external data, enabling the model to better respect the intended trust hierarchy. Experimental results show that LTBD substantially outperforms inference-time defenses and performs competitively with training-based approaches, while preserving benign-task utility and introducing negligible inference overhead. In particular, LTBD achieves 0.00% ASR on AlpacaFarm and only 0.11-0.19% ASR on TaskTracker. LTBD also remains effective under adaptive attacks, where adversaries have full knowledge of the defense and explicitly attempt to bypass it.

补充信息

↑