arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AKRASIA:针对基于推理的代码大语言模型的隐蔽后门攻击

AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs

Chua Jin Chou, Sarang Nambiar, Murali Srinivasan, Ezekiel Soremekun

arXiv 2609.01023首次发表:更新:

发表机构

Singapore University of Technology and Design; International Institute of Information Technology Bangalore(新加坡科技设计大学; 班加罗尔国际信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员提出AKRASIA隐蔽后门攻击,针对基于推理的代码大语言模型,可实现高攻击成功率且能规避防御与人工检查,凸显了防御大语言模型推理后门的必要性。

AI 中文摘要

我们提出了AKRASIA,这是一种针对基于推理的代码大语言模型(Code LLMs)的隐蔽推理时后门攻击。AKRASIA旨在在推理大语言模型中实现后门目标(如恶意代码执行),同时规避自动防御和人工检查。为实现这一目标,AKRASIA探测目标大语言模型以构建代码级后门触发器,随后采用上下文学习进行后门学习,并利用模型不忠实性来隐藏后门触发器,生成合理的推理过程。我们使用4个后门目标、6个推理大语言模型、3个编码任务/数据集和3种防御方法对AKRASIA进行评估。AKRASIA在最先进(SOTA)大语言模型上的平均攻击成功率高达99.34%,同时保持高达97.23%的平均准确率。AKRASIA规避了最先进的防御,在大多数(14/18)防御设置中保留了高达98.82%的平均攻击成功率,还规避了人工检查,在多达80%的设置中成功隐藏了后门触发器和推理步骤。我们的发现凸显了防御大语言模型免受推理后门攻击的必要性。

英文摘要

We present AKRASIA, a stealthy, inference-time backdoor attack against reasoning-based Code LLMs. AKRASIA aims to achieve a backdoor target (e.g., malicious code execution) in reasoning LLMs while evading automated defenses and human inspection. To achieve this, AKRASIA probes the victim LLM to construct a code-level backdoor trigger. It then employs in-context learning for backdoor learning, and model unfaithfulness to conceal the backdoor trigger, and generate plausible reasoning. We evaluate AKRASIA using four backdoor targets six (6) reasoning LLMs, three coding tasks/datasets and three defense methods. AKRASIA has up to 99.34% average attack success rate on SOTA LLMs and mantains up to 97.23% average accuracy. AKRASIA evades the SOTA defense, retaining up to 98.82% average ASR in most (14/18) defense settings. It evades human inspection, successfully hiding the backdoor trigger and reasoning steps in up to 80% of settings. Our findings motivate the need to defend LLMs against reasoning backdoors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑