arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02358cs.CL

ScrambleToolBench:即使自身的“地图”指向下一步,智能体仍会进行穷尽式搜索

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有工具使用基准依赖先验知识的局限,推出ScrambleToolBench交互式终端基准,评估发现当前语言模型面对环境结构变化时难以实现稳健适应。

中文摘要 AI 辅助

为了在开放世界环境中稳健运行,自主智能体应仅通过交互就能推断陌生系统的行为,即便没有文档支持。然而,现有的工具使用基准在静态环境中暴露语义工具模式,使智能体能够依赖先验知识而非自主发现。为解决这一局限,我们推出ScrambleToolBench,这是一个交互式终端基准,旨在隔离行为推理。通过移除语义线索并实施连续任务课程,该基准要求智能体完全通过试错交互来揭示隐藏的工具行为。该基准还引入了动态挑战,包括地图漂移、随机动作失败和时间执行窗口,以评估智能体是否能在环境变化时修正并调整其假设。我们对最先进的语言模型的评估显示,成功的初始发现并不能转化为稳健的适应能力。当面临地图漂移等结构变化时,智能体无法使用循环追踪等演绎策略,反而表现出信念惯性或退回到穷尽式搜索。增加测试时推理只会放大这种昂贵的蛮力搜索,而无法实现演绎式恢复。虽然为智能体配备持久内存可减少复合错误,但它们仍然无法有效推断结构变化,凸显了当前智能体推理存在的缺口。

英文摘要

To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, existing tool-use benchmarks expose semantic tool schemas in static environments, allowing agents to rely on prior knowledge rather than autonomous discovery. To address this limitation, we introduce ScrambleToolBench, an interactive terminal benchmark designed to isolate behavioral reasoning. By removing semantic cues and enforcing a continuous task curriculum, the benchmark requires agents to uncover hidden tool behaviors entirely through trial-and-error interaction. The benchmark further introduces dynamic challenges, including mapping drift, stochastic action failures, and temporal execution windows, to evaluate whether agents can revise and adapt their hypotheses as the environment changes. Our evaluation of state-of-the-art language models reveals that successful initial discovery does not translate into robust adaptation. When faced with structural changes such as mapping drift, agents fail to use deductive strategies such as cycle tracing, and instead exhibit belief inertia or fall back to exhaustive search. Increasing test-time reasoning only amplifies this expensive brute-force search rather than enabling deductive recovery. While equipping agents with persistent memory reduces compounding errors, they remain unable to efficiently infer structural changes, highlighting a gap in current agent reasoning.

发表机构

  • Nanyang Technological University(南洋理工大学)
  • Agency for Science, Technology, and Research (A*STAR)(新加坡科学、技术与研究局)

机构由 AI 辅助整理,请以论文原文为准。

↑