基于置换的大型语言模型隐恶意软件:威胁与对策
Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures
浏览论文内容
中文总结 AI 辅助
本文利用行为保持置换对称性,既作为中和LLM隐恶意软件的防御手段,也作为理论上无损的攻击编码方法,并量化了其性能损失极小。
中文摘要 AI 辅助
训练大型语言模型(LLM)的难度及其普遍性,加剧了隐恶意软件(stegomalware)的威胁,即恶意载荷被嵌入到模型权重中。近期工作已展示利用模型权重中的置换对称性来缓解这些威胁,但未能证明对所有LLM权重中的隐恶意软件进行中和。在本文中,我们展示了行为保持对称性作为抵御隐恶意软件防御手段的全部潜力,以及这些对称性在被攻击者利用时所带来的风险。对于隐恶意软件中和,我们改进了先前的工作,证明可以选择置换来移动所有模型参数。这与先前方法形成对比,先前方法在LLM中留下了相当大比例的权重未改变。当用于攻击时,我们表明置换对称性可以将恶意软件以理论上无损的方式编码到模型权重中,编码后无需重新训练,且提取脚本中不需要载荷特定信息——这一组合特性在以往任何单一方法中均未出现过。虽然理论上无损,但置换在实践中可能因数值误差的累积而改变模型行为。因此,我们量化了应用这些方法(无论用于攻击还是防御)所导致的模型性能损失,结果表明损失极小。
英文摘要
The difficulty of training large language models (LLMs), together with their ubiquity, raises the threat of stegomalware, where malicious payloads are embedded into model weights. Recent work has demonstrated the use of permutation symmetry in model weights to mitigate these threats, but failed to show neutralization of stegomalware across all weights for LLMs. In this paper, we demonstrate the full potential of behavior-preserving symmetries as a defense against stegomalware, as well as the risks these symmetries pose when exploited by attackers. For stegomalware neutralization, we improve upon previous work, demonstrating that it is possible to select permutations which displace all model parameters. This contrasts with previous methods which left a significant percentage of weights unaltered in LLMs. When used in an attack, we show that permutation symmetries can encode malware into the weights of a model in a way that is theoretically lossless, requires no retraining after encoding, and needs no payload-specific information in the extraction script---a combination of characteristics not previously seen in any single method. While theoretically lossless, permutation can in practice alter model behavior due to the accumulation of numerical error. We therefore quantify the loss in model performance associated with applying these methods, for both attack and defense, showing it to be minimal.
发表机构
- Fuzzy Labs(模糊实验室)
机构由 AI 辅助整理,请以论文原文为准。