AppDeltaWorld:面向移动GUI智能体的基于转换的增量代码世界模型
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
浏览论文内容
中文总结 AI 辅助
针对移动GUI智能体轨迹获取难、模拟环境扩展难等问题,提出AppDeltaWorld增量代码世界模型,在CMGUIBench-500评估中表现最优,支撑的AppDeltaAgent在多基准达SOTA,测试时强化学习可进一步提升性能。
中文摘要 AI 辅助
移动GUI智能体可通过像素感知和触摸操作来操作应用,是收集和改进长程移动交互策略的有前景的接口。然而,敏感应用和隐私关键操作难以获取真实轨迹;同时,现有模拟环境难以扩展,GUI世界模型仍存在生成不稳定、模态覆盖有限、操作转换逻辑不一致的问题。为解决这些局限,我们提出AppDeltaWorld,一种基于转换的增量代码世界模型,它将下一GUI预测为可达的代码更新,而非无约束的图像或文本描述。AppDeltaWorld在操作转换约束下检索应用特定的一级HTML参考,基于当前屏幕、操作、预测的下一屏幕文本和检索到的结构生成二级可执行HTML,并在浏览器渲染前将生成的视觉资产插入图像槽。作为世界模型,AppDeltaWorld在CMGUIBench-500的Code2World评估中实现最高保真度,在结构布局和UI元素重建上较仅图像和仅代码基线有明显提升;作为训练环境,AppDeltaWorld支持过滤后的闭环SFT数据构建,结合公共监督使AppDeltaAgent在AndroidLens上实现SOTA性能,在MobileGym和MobileWorld上实现一致提升;此外,基于世界模型的测试时强化学习支持策略适应,无需与真实应用额外交互即可进一步改进。
英文摘要
Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies. However, real trajectories are difficult to obtain for sensitive apps and privacy-critical operations. At the same time, existing simulated environments are costly to scale up, and GUI world models still suffer from unstable generation, limited modality coverage, and inconsistent action-transition logic. To address these limitations, we propose AppDeltaWorld, a transition-grounded delta code world model that predicts the next GUI as a reachable code update rather than as an unconstrained image or text description. AppDeltaWorld retrieves app-specific Level-1 HTML references under an action-transition constraint, generates Level-2 executable HTML conditioned on the current screen, action, predicted next-screen text, and retrieved structure, and inserts generated visual assets into image slots before browser rendering. As a world model, AppDeltaWorld achieves the highest fidelity on CMGUIBench-500 under Code2World evaluation, with clear gains in structural layout and UI element reconstruction over image-only and code-only baselines. As a training environment, AppDeltaWorld supports filtered closed-loop SFT data construction that, when combined with public supervision, enables AppDeltaAgent to achieve state-of-the-art performance on AndroidLens and consistent gains on MobileGym and MobileWorld. Moreover, world-model-based test-time reinforcement learning enables policy adaptation and shows further improvements without additional interaction with real apps.