arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语用攻击面:大语言模型中隐式上下文的漏洞

Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

Bocheng Chen, Han Zi, Roucheng Ou, Yawei Liu, Minyue Chen, Zimo Qi, Rongrong Wang, Guangliang Liu

arXiv 2608.09551首次发表:更新:

发表机构

University of Mississippi; Michigan State University; Johns Hopkins University; Indiana University Indianapolis(密西西比大学; 密歇根州立大学; 约翰斯·霍普金斯大学; 印第安纳大学印第安纳波利斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出大语言模型存在语用攻击面漏洞,利用人类语言解读与安全对齐的隐式上下文不匹配,所提攻击方法在各类模型上的成功率大幅优于基线方法。

AI 中文摘要

在大语言模型(LLM)时代,攻击者常操控自然语言诱导出不安全或有害的输出,形成了基于LLM的系统特有的新型自然语言攻击面,此类攻击直接利用用户提示中的显式语言线索来绕过LLM的安全机制,不过现有安全对齐算法通常可缓解这类攻击。另一方面,人类语言本质上基于语用学,需借助典型上下文(如世界知识、社会规范)来解读语言,但这类上下文往往是隐式的,因为它们未直接在人类语言中表达,且在安全对齐中未被充分利用,造成了人类语言解读与安全对齐方法之间的根本不匹配。本文证明这种不匹配会暴露LLM的漏洞,将该漏洞称为语用攻击面,可被利用以实现高攻击成功率。实验结果表明,所提方法在各类开源和闭源模型上的攻击成功率大幅优于基线攻击方法。

英文摘要

In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique to LLM-based systems, where attacks directly exploit explicit linguistic cues in user prompts to bypass the safety mechanism of LLMs. However, such attacks can often be mitigated by existing safety alignment algorithms. On the other hand, human language is inherently grounded in pragmatics, necessitating typical context to interpret language, e.g., world knowledge, social norms. However, such contexts are often implicit because they are not directly expressed in human language and are not sufficiently leveraged in safety alignment, creating a fundamental mismatch between human language interpretation and safety alignment approaches. In this paper, we demonstrate that this mismatch exposes vulnerabilities in LLMs. We refer to this vulnerability as the pragmatic attack surface, which can be exploited to achieve high attack success rates. The experimental results demonstrate that our proposed approach outperforms baseline attack methods across various open-source and closed-source models by a substantial margin.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑