审计大语言模型智能体动作选择中的溯源敏感性
Auditing Provenance Sensitivity in LLM Agent Action Selection
浏览论文内容
中文总结 AI 辅助
研究大语言模型智能体动作选择中的溯源敏感性,引入针对特定目标的授权审计,通过改变命题源权威等测试行为,发现可信和不可信变体在部分案例中产生不同动作,模型对线索有反应但不可信证据仍影响行动。
中文摘要 AI 辅助
大语言模型智能体从包含用户请求、工具输出、检索记录、内存和不可信文本的上下文中选择工具和论据。证据可能相关但未经授权决定决策,因此正确行动不必仅基于允许的证据。我们引入了一种针对特定目标的授权审计,为每个工具和论据目标分别标记上下文因素。主要测试在保持任务、命题、立场和政策不变的情况下,仅改变命题的源权威。然后在有效证据被削弱时测试行为,并使用上下文子集交互作为辅助定位诊断。在450个受控的下一个动作任务和多个开放权重的大语言模型系列中,可信和不可信变体在5.4%的竞争案例中产生不同动作,而在支持案例中为1.7%。在受控降级下,2.4%的比较中未经授权的竞争以完全正确、混合错误、完全正确的模式保留,95%置信区间为2.1%至3.0%。这些是受控压力集比率,而非部署流行率。模型对文本源权威线索有反应,但这并不能阻止不可信证据影响其行动。
英文摘要
LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action need not be grounded only in permitted evidence. We introduce a target-specific authorization audit that labels context factors separately for each tool and argument target. Its primary test holds the task, proposition, position, and policy fixed while changing only the proposition's source authority. We then test behavior when valid evidence is weakened and use context-subset interactions as a secondary localization diagnostic. Across 450 controlled next-action tasks and multiple open-weight LLM families, trusted and untrusted variants produce different actions in 5.4 percent of competing cases versus 1.7 percent of supporting cases. Under controlled degradation, unauthorized competition is retained in a full-correct, mixed-error, clean-correct pattern in 2.4 percent of comparisons, with a 95 percent confidence interval from 2.1 to 3.0 percent. These are controlled stress-set rates, not deployment prevalence. The models respond to textual source-authority cues, but this does not prevent untrusted evidence from influencing their actions.