arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于低秩防御与电路引导代理的高效大语言模型对抗训练

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates

Weiyi He, Yuping Lin, Jiliang Tang, Yue Xing

arXiv 2607.28959首次发表:更新:

发表机构

Michigan State University(密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对大语言模型对抗训练计算成本高的问题,从防御侧优化表示微调、攻击侧构建轻量代理模型两方面提出策略,使每步对抗训练FLOPs降48.1%,仅需0.0118%可训练参数。

AI 中文摘要

对抗训练是抵御对抗攻击最有效的防御手段之一,但在现代规模下,尤其是针对大语言模型(LLM)时,其计算成本仍高得令人却步。尽管已有缓解策略如潜在对抗训练(LAT)被提出,但它们仍会产生高昂的计算成本。本研究从两个互补视角全面探究加速LAT的高效计算策略:(1)防御侧优化:在LAT中探究表示微调(ReFT),并揭示若对标记应用ReFT与攻击的对象不匹配时存在潜在问题;(2)攻击侧优化:在每次LAT迭代中计算对抗攻击时,仅从LLM中提取相关电路以构建轻量代理模型,避免攻击生成时通过完整模型进行前向-反向传播的计算。针对两个视角,本研究均提供理论依据与数值证据以说明所提策略的有效性。最终,与采用全微调的标准LAT相比,本方法平均将每步对抗训练的浮点运算量(FLOPs)降低48.1%,且仅需0.0118%的可训练参数。

英文摘要

Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). While existing mitigation strategies, e.g., latent adversarial training (LAT), have been developed, they still incur a high computational cost. In this work, focusing on LLM-based classifiers, we investigate computation-efficient strategies to speed up LAT from two complementary perspectives: (1) Defense-side optimization: We explore representation fine-tuning (ReFT) within LAT, and reveal a potential issue if there is a mismatch on which tokens to apply ReFT and the attack. (2) Attack-side optimization: When computing adversarial attacks in each LAT iteration, we extract only the relevant circuits from the LLM to construct a lightweight surrogate model, avoiding the computation in the forward-backward passes through the full model during the attack generation. For both perspectives, we provide theoretical justifications and numerical evidence to illustrate the effectiveness of the proposed strategies. Compared to standard LAT with full fine-tuning, our method on average reduces per-step adversarial-training FLOPs by 29.6% while requiring only 0.0118% trainable parameters, at a moderate cost in robustness.

Comments30 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑