arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00048cs.CLcs.AI

GUI-CC:将GUI世界模型作为智能体环境对其上下文一致性进行基准测试

GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments

  • Yale University(耶鲁大学)
  • Zhejiang University(浙江大学)
  • China University of Geosciences(中国地质大学)
  • DAMO Academy, Alibaba Group(阿里巴巴达摩院)
  • Tongji University(同济大学)
  • University of California, San Diego(加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

Lin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu, Yilun Zhao, Yu Rong

AI总结:

研究针对GUI世界模型作为智能体环境评估的不匹配问题,提出GUI-CC基准,含两个赛道及对应任务,评估四项指标,发现当前模型单步生成合理但多步上下文一致性不足。

AI中文摘要:

GUI世界模型越来越多地被作为单步下一个屏幕预测器进行评估,但其预期用途通常是作为GUI智能体的多步环境。这种不匹配使得一项关键要求未得到充分测试:当生成的状态被重复用于未来交互时,必须保持上下文一致性。我们提出GUI-CC,这一基准将GUI世界模型作为智能体环境而非孤立的下一个屏幕预测器来评估其上下文一致性。GUI-CC包含两个互补的赛道:离线参考动作赛道,该赛道沿真实移动GUI轨迹滚动模型;以及在线智能体循环赛道,该赛道让固定探测智能体与模型生成的UI进行交互。我们从GUIOdyssey构建了500个离线轨迹任务,并在30个移动应用程序中构建了200个经模拟器验证的在线任务。GUI-CC评估转换保真度、转换合理性、上下文一致性和任务进度。实验表明,合理的单步生成并不能保证可靠的环境模拟:当前模型常常生成看起来可用的屏幕,却未能保留与任务相关的上下文,也无法支持可执行的多步滚动。

英文摘要:

GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use is often as multi-step environments for GUI agents. This mismatch leaves a key requirement under-tested: generated states must remain contextually consistent when they are repeatedly reused for future interaction. We introduce GUI-CC, a benchmark that evaluates contextual consistency of GUI world models as agent environments rather than isolated next-screen predictors. GUI-CC contains two complementary tracks: an offline reference-action track that rolls models along real mobile GUI trajectories, and an online agent-loop track that lets fixed probing agents interact with model-generated UIs. We construct 500 offline trajectory tasks from GUIOdyssey and 200 emulator-verified online tasks across 30 mobile apps. GUI-CC evaluates transition fidelity, transition plausibility, contextual consistency, and task progress. Experiments show that plausible single-step generation does not guarantee reliable environment simulation: current models often produce usable-looking screens while failing to preserve task-relevant context or support executable multi-step rollouts.

补充信息

↑