AI 中文总结
研究当可靠工具在会话中悄然改变时语言模型智能体的工具选择,借鉴认知心理学的集合转移构建基准和评估框架,计算集合转移准确率,测试发现不同失败模式及集合框架对路由动态有影响。
AI 中文摘要
当可靠工具在正在进行的会话中悄然改变时,语言模型智能体的工具选择会发生什么?我们借鉴认知心理学中的集合转移来研究智能体如何适应隐藏的可靠性变化。我们的基准构建了具有冗余的工具技能库,其中许多工具解决相同任务但隐藏可靠性不同。在我们的评估框架中,分支调度在隐藏边界处转移可靠工具组,并将每次转移与无转移控制配对。我们发现,默认情况下,智能体在每个边界的几轮内就会固定在一个小的重复例程上,每次可靠性转移后调用份额集中在几个离散值上。我们为每个智能体轨迹计算集合转移准确率:在每个转移后窗口中路由到目标工具组的联合概率。我们在开源的智能体框架中测试开放权重的语言模型,发现在同一组例程中存在质的不同的失败模式。我们还发现,集合框架,即工具集将替代方案呈现为竞争或互补的方式,会改变路由动态。
英文摘要
What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow the notion of set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our cognitive test for LLM agents mounts libraries of redundant tools and skills, in which many tools solve the same task but differ in hidden reliability. Using a branching schedule, we shift the reliable tool group in the environment and compare it with a stable control, allowing us to isolate the effect of each shift on the agent's behavior. We conduct our study on a panel of LLMs equipped with harnesses and show that the same set of shifts results in distinct behaviors across models: some latch onto a fixed routine within a few turns, whereas others continue to vary. Less capable models often omit the reliable tool group, while frontier models keep calling it alongside the other groups. We introduce a suite of measures to quantify agent behavior after reliability shifts. While policy prompting substantially alters behavior in some tested models, our findings highlight agents' brittleness when changes occur in indirectly observable context.
CommentsAccepted at COLM 2026 Workshop on Agent Behavior