发表机构
Information Engineering University; College of Computer Science and Technology, Zhejiang University; State Key Laboratory of Information Security, Institute of Information Engineering, Chinese Academy of Sciences(信息工程大学; 浙江大学计算机科学与技术学院; 中国科学院信息工程研究所信息安全国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型在高风险决策场景中的决策级劫持问题,提出CogBias框架,利用位翻转攻击,通过可微情感评估器等将主观偏好转化为优化信号,经实验验证该方法能在少量位翻转下诱导目标主题立场转变,破坏模型价值对齐。
AI 中文摘要
大语言模型(LLMs)已广泛应用于企业战略等高风险决策场景,用户越来越依赖其输出。开源模型共享生态系统与基于LLM的关键决策应用的深度整合带来了关键风险,即攻击者若能操纵模型认知立场,就能间接影响下游决策者的判断和行动,本文将此类威胁定义为决策级劫持。现有攻击无法在不触发违禁内容或不降低模型功能的情况下实现有针对性的认知操纵。本文揭示位翻转攻击(BFAs)可作为诱导决策级劫持的攻击向量,部署后只需翻转少量权重位就能实现隐蔽、低成本且持久的认知操纵。因此提出了CogBias,即一种针对LLMs的认知偏差注入框架。CogBias通过可微情感评估器将主观偏好转化为优化信号,使用多目标损失联合约束多个维度,并构建BitScout定位关键位,在超稀疏翻转预算下实现有针对性的认知干预。在Llama-3.2-3B、Mistral-7B和Qwen2.5-14B以及商业推荐和有争议的事实性话题场景上的实验表明,仅翻转少量位就能在目标主题上稳定地诱导显著的立场转变,而对非目标任务和整体输出分布的影响有限。这项工作表明对低级权重数据的微小扰动足以破坏LLMs的高级价值对齐。
英文摘要
Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly relying on their outputs. However, the deep integration of open-source model sharing ecosystems with LLM-powered critical decision-making applications also introduces critical risks: if an attacker can manipulate the model's cognitive stance, they can indirectly influence the judgments and actions of downstream decision-makers. This paper defines such threats as decision-level hijacking. Existing attacks fail to achieve targeted cognitive manipulation without triggering prohibited content or degrading model functionality. To fill this gap, this paper reveals that Bit-Flip Attacks (BFAs) can serve as an attack vector for inducing decision-level hijacking, requiring no real-time interaction or control over the training process, and only a minimal number of weight bits need to be flipped after deployment to achieve stealthy, low-cost, and persistent cognitive manipulation. Therefore, we propose CogBias, a cognitive bias injection framework for LLMs. CogBias converts subjective preferences into optimization signals via a differentiable sentiment evaluator, uses a multi-objective loss to jointly constrain multiple dimensions, and constructs BitScout to locate critical bits, achieving targeted cognitive intervention under an ultra-sparse flip budget. Experiments on Llama-3.2-3B, Mistral-7B, and Qwen2.5-14B, as well as on the commercial recommendation and controversial factual topic scenarios, demonstrate that flipping only a small number of bits stably induces significant stance shifts on target topics, while the impact on non-target tasks and overall output distribution is limited. This work demonstrates that minute perturbations to low-level weight data suffice to undermine the high-level value alignment of LLMs.