TrustMI:因果控制助手如何信任其用户
TrustMI: Causally controlling how assistants trust their users
- Inria Paris(巴黎国立信息与自动化研究所)
- Sorbonne Université(索邦大学)
- Meta SuperIntelligence Labs(Meta超级智能实验室)
- LightOn(LightOn公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出TrustMI方法,通过构建对比对话并学习激活转向矩阵,在冻结参数下因果控制LLM助手对用户的信任决策,并在多模型及安全场景中验证其单调效应。
AI中文摘要:
大型语言模型(LLM)助手通常会决定是否信任那些能力、意图和诚信无法验证的用户及第三方。这种不确定性对安全性至关重要,因为信任错误的方可能导致智能体遵从有害请求,或在工具使用过程中执行恶意指令。为研究该问题,我们将信任定义为助手愿意接受因另一方行为而带来的脆弱性,并探究此类行为是否可以通过模型激活进行因果控制。我们构建了2,000个涵盖能力、善意和诚信维度的对比对话,其中配对响应完成相同请求,但区别在于助手是否信任用户。基于这些配对,我们在保持模型参数冻结的情况下学习转向矩阵,并在来自三个家族的六个指令微调模型上进行测试,发现转向能在两个方向上单调地改变信任决策。随后,我们考察该效应是否扩展到涉及有害请求、提示注入和内部威胁的若干安全相关智能体场景,并以良性任务和推理作为对照。我们的研究结果证明,对用户的信任可以沿模型激活中的线性方向进行因果控制,并提供了一种研究信任如何塑造语言模型中安全相关行为的方法。
英文摘要:
Large Language Model (LLM) assistants routinely decide whether they can trust users and third parties whose competence, intentions, and integrity they cannot verify. This uncertainty matters for safety, as trusting the wrong party can lead an agent to comply with harmful requests or act on malicious instructions encountered during tool use. To study this problem, we define trust as an assistant's willingness to accept vulnerability to the actions of another party and ask whether such behavior can be causally controlled through model activations. We build 2,000 contrastive conversations spanning ability, benevolence, and integrity, where paired responses complete the same request but differ in whether the assistant trusts the user. From these pairs, we learn steering matrices while keeping the model parameters frozen and test them across six instruction-tuned models from three families, finding that steering changes trust decisions monotonically in both directions. We then ask whether this effect extends to several safety-related agent settings involving harmful requests, prompt injections, and insider threats, while using benign-task and reasoning as controls. Our findings provide evidence that trust in the user can be causally controlled along linear directions in model activations and provide a way to study how trust shapes safety-relevant behavior in language models.