arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34769cs.CLcs.AI

LongPuzzleBench:评估GUI智能体在长时程视觉谜题上的表现

LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles

Bingo Zhang, Haochuan Lu, Zongjie Li, Genjian Li, Ari Yu Zhang, Chaozheng Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出LongPuzzleBench基准,通过六个谜题游戏的114个关卡测试GUI智能体长时程视觉推理,发现智能体因仅依据可见进展决策而失败,人类可解决所有目标。

中文摘要 AI 辅助

GUI智能体需要长时程的视觉推理能力:它们必须解读不断变化的界面,同时保持多步骤计划的可行性,因为先前的行动会限制后续的行动。现有的基准测试评估了基础操作、计算机使用和游戏玩法,但很少测试智能体在长链耦合决策中是否保持连贯性。长时程视觉谜题直接暴露了这一能力:一个看似有进展的合法移动可能使谜题变得无解,而损失要到几步之后才会显现。我们推出了LongPuzzleBench,包含六个谜题游戏中的114个关卡,通过原生GUI操作进行,其中一个目标可能需要人类在持久棋盘上执行超过一千次操作,且死胡同不会提前宣告。仅使用原生GUI操作时,最强的智能体解决了大多数目标,但在更难、更长的棋盘上成功率急剧下降:十个通用智能体中有七个无法解决比中等难度更难的问题,且没有一个能完成Bolt Unscrew Hard,而人类可以解决该关卡以及所有其他目标。代码执行CUA并未缩小这一差距,其得分混合了视觉求解与算法搜索。受控诊断将这些失败追溯到单一局限,且该局限无法通过规则、状态提示或失败记忆来消除:智能体根据每一步可见的进展来判断移动,而非根据其留下的未来选项。

英文摘要

GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains of coupled decisions. Long-horizon visual puzzles expose this capability directly: a legal move that looks like progress can make the puzzle unsolvable, and the loss shows only several moves later. We introduce LongPuzzleBench, 114 levels in six puzzle games played through native GUI actions, where one objective can take a human over a thousand actions on persistent boards and dead ends go unannounced. With Native GUI Actions alone, the strongest agents solve most objectives, but success falls sharply on harder, longer boards: seven of ten general-purpose agents solve nothing harder than Medium, and none completes Bolt Unscrew Hard, which a human solves along with every other objective. Code Execution CUA does not close this gap, and its scores mix visual solving with algorithmic search. Controlled diagnostics trace these failures to one limitation that neither rules, state hints, nor failure memory removes: agents judge each move by the visible progress it makes, not by the future options it leaves.

发表机构

  • Vera Praxis
  • Tencent(腾讯)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑