arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

技能使用还是技能表演?评估技能增强型语言智能体中的推理后台

Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents

Jinwei Hu, Yi Qi, Xinmiao Huang, Youcheng Sun, Yi Dong, Xiaowei Huang

arXiv 2607.27484首次发表:更新:

AI 中文总结

本研究针对技能增强型语言智能体,提出BACKTRACE评估框架及BACKROOMBench测试平台,发现其存在推理后台现象,即技能使用声明与实际决策影响存在系统性差距,需通过干预进行审计。

AI 中文摘要

可复用技能正成为为语言智能体扩展任务流程的标准接口。然而评估者通常从可见推理或智能体自身的归因推断技能使用,这些信号显示智能体看似使用的内容,而非技能是否改变了其决策。本文探究技能增强型智能体是否存在“推理后台”,即所声明的技能使用与干预测量出的影响之间的系统性差距。我们提出BACKTRACE评估框架,将每个基于技能的答案与匹配的无技能反事实结果配对,对技能的含义、措辞、身份、内容及分配进行干预,并在答案确定后才引出归因。我们将该框架实例化为BACKROOMBench,这是一个经过验证的测试平台,涵盖受控逻辑与竞赛数学、多种技能条件、单智能体和多智能体设置以及不同模型系列。我们的评估揭示了普遍的来源失败:在所有模型和领域中,所声明的技能使用往往保持稳定,而因果依赖和符号效用却发生变化,产生隐性采用和表演性使用。行为效果对程序内容的响应比对显示的技能身份更可靠,而所声明的归因对人工制品可用性的响应强烈。基于直接技能使用声明、文本提及、轨迹相似性和LLM评判的观测检测器无法识别哪些决策实际依赖于技能。在多智能体系统中,技能影响即使在其来源丢失后仍可在通信中保留,而无技能团队仍会命名从未提供的技能和来源。这些发现确立了推理后台是一个通用的AI来源问题,其审计需要干预。

英文摘要

Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's own attribution. These signals show what the agent appears to use, not whether the skill changed its decision. We ask whether skill-augmented agents exhibit a \textbf{Reasoning Backroom}, a systematic gap between stated skill use and intervention-measured influence. We introduce BACKTRACE, an evaluation framework that pairs each skill-conditioned answer with a matched no-skill counterfactual, intervenes on skill meaning, wording, identity, content, and assignment, and elicits attribution only after the answer is committed. We instantiate the framework as BACKROOMBench, a verified testbed spanning controlled logic and competition mathematics, multiple skill conditions, single-agent and multi-agent settings, and diverse model families. Our evaluation reveals a pervasive provenance failure. Across models and domains, stated skill use often remains stable while causal reliance and signed utility vary, producing both silent uptake and performative use. Behavioral effects follow procedural content more reliably than displayed skill identity, whereas stated attributions respond strongly to artifact availability. Observational detectors based on direct skill-use claims, text mentions, trace similarity, and an LLM judge do not identify which decisions actually depend on the skill. In multi-agent systems, skill influence can survive communication even after its source is lost, while no-skill teams still name skills and sources that were never supplied. These findings establish the Reasoning Backroom as a general AI provenance problem whose audit requires intervention.

Comments21 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑