arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估与改进大型语言模型对输入序列变异的鲁棒性

Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

Narek Maloyan

arXiv 2610.02432首次发表:更新:

AI 中文总结

针对LLM面临的输入变异攻击,提出鲁棒性度量R_stab、攻击方法ASA及防御策略(如AttestMCP),显著降低攻击成功率,提升系统安全性。

AI 中文摘要

生产系统中的大型语言模型(LLM)面临提示注入、木马(后门)以及自动质量指标的操纵。本论文开发了用于评估和改进LLM对对抗性输入序列变异鲁棒性的模型、方法和算法。我们提出了R_stab(f),一种基于小输入扰动下每步输出分布之间的Jensen-Shannon散度的生成式鲁棒性度量。对于局部攻击,我们证明了V(h) <= 1 - R_class(h),其中R_class(h)是决策算子h在小扰动下保持其决策的概率。对于非局部攻击,我们提出了一种校准的经验模型。对于LLM作为评判者的系统,我们开发了ASA,一种自适应进化黑盒攻击,其攻击成功率(ASR)最高可达73.8%,在开放模型之间的迁移率高达62.6%。在Trojan Detection Challenge 2023数据(Pythia-1.4B)上,替代触发器达到REASR约0.99,而真实触发器的召回率约为0.17,基线约为0.14。在SaTML CTF 2024上,我们系统化了四类绕过多层防御的方法,将ASR从90%降低到15-25%。由5-7个异构模型组成的委员会将Gemma-3-4B的ASR降低了47-55个百分点,使用7个模型时降至19.3%。对于基于模型上下文协议(MCP)的智能体系统,我们提出了AttestMCP,它使用HMAC保护的数据包对工具调用进行认证,每次调用耗时低于0.1毫秒,以及提交边界隔离模式。在包含847个场景的MCPBench基准上,它们将平均ASR从53.7%降至12.4%。这些方法已在JudgeGuard和TrojanArmor软件套件以及MCPSec模块中实现。

英文摘要

Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.

CommentsPhD thesis, 2026. 118 pages. v2: corrected a name spelling in the acknowledgements

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑