arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AGENTQ:面向LLM智能体的量化条件后门攻击

AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents

Xiaoqun Liu, Qiben Yan

arXiv 2609.14060首次发表:更新:

发表机构

Michigan State University(密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM智能体的量化条件后门攻击,提出AGENTQ框架,结合分层LoRA注入与部分PGD修复,在保持良性效用下实现高达100%的量化后攻击成功率。

AI 中文摘要

量化是开放权重LLM智能体的默认部署路径之一,但它并非保持行为不变:攻击者可以发布一个通过审计的全精度检查点,但该检查点在量化后会出现异常行为,这被称为量化条件攻击(QCA)。先前的QCA工作针对自由文本生成,其危害由人类读者介导。相比之下,智能体环境构成更严重的风险:触发的载荷是结构化函数,可以在无人监督的情况下执行。我们首次对针对LLM智能体的QCA进行研究。我们发现,直接改编先前的后门注入方法可以在量化后产生恶意行为,但会显著降低良性效用,使所得攻击不切实际。为了理解威胁的真正上限,我们提出AGENTQ,一种攻击框架,将分层带LoRA注入与多码本量化等价类上的部分PGD修复相结合。AGENTQ在保持正常智能体能力的同时,将恶意行为集中在量化模型中。在三个触发-动作对和三个码本(NF4、FP4、INT8)上,AGENTQ达到高达100%的量化后攻击成功率,且良性效用损失极小,这强调了在部署开放权重智能体之前,将量化感知安全评估作为标准要求的必要性。

英文摘要

Quantization is one of the default deployment paths for open-weight LLM agents, but it is not behavior-preserving: an adversary can release a full-precision checkpoint that passes audits yet misbehaves once quantized, termed as quantization-conditioned attack (QCA). Prior QCA work targets free-text generation, where harm is mediated by a human reader. In contrast, the agentic setting poses a more severe risk: the triggered payload is a structured function that can be executed without human oversight. We present the first study of QCA against LLM agents. We find that directly adapting prior backdoor-injection methods can produce malicious behavior after quantization, but substantially degrades benign utility, rendering the resulting attacks impractical. To understand the true upper bound of the threat, we propose AGENTQ, an attack framework that combines layer-banded LoRA injection with partial-PGD repair over a multi-codebook quantization-equivalence class. AGENTQ preserves normal agentic capability while concentrating malicious behavior in the quantized model. Across three trigger-action pairs and three codebooks (NF4, FP4, INT8), AGENTQ reaches up to 100% post-quantization attack success rate with minimal loss of benign utility, underscoring the need to make quantization-aware safety evaluation a standard requirement before open-weight agents are deployed.

CommentsAppear in EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑