arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13287cs.CVcs.AI

LLaDA-UI:将块级扩散引入视觉语言GUI智能体

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

发表机构Inclusion AI · 西湖大学
查看机构详情
  • Inclusion AI
  • Westlake University(西湖大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhangxuan Gu, Haoxing Chen, Qi Qin, Yi Xin, Kai Gan, Lin Liu, Long Cui, Xiaomei Wang, Beitong Zhou, Yunzhu Zhang, Zhengwen Zeng, Changlong Gao, Weizhi Chen, Ron… 展开作者

Zhangxuan Gu, Haoxing Chen, Qi Qin, Yi Xin, Kai Gan, Lin Liu, Long Cui, Xiaomei Wang, Beitong Zhou, Yunzhu Zhang, Zhengwen Zeng, Changlong Gao, Weizhi Chen, Rongchao Zhang, Haoyuan Wu, Shuheng Shen, Changhua Meng, Weiqiang Wang, Jianguo Li, Zhenzhong Lan

首次发表
浏览论文内容

中文总结 AI 辅助

LLaDA-UI提出基于MoE的167亿参数块级扩散视觉语言GUI智能体,通过两阶段训练在多个GUI基准上超越Qwen2.5-VL-7B和Qwen3-VL-8B,验证了块级扩散的实用性。

中文摘要 AI 辅助

扩散大语言模型(dLLMs)通过块并行、任意顺序生成实现了高解码效率,使其对延迟敏感的应用具有吸引力。GUI智能体是这一范式的自然测试平台,因为它们必须反复感知屏幕状态并实时发出结构化、空间定位的动作。然而,dLLMs能否在保持其并行解码优势的同时扩展为多模态GUI智能体仍是一个开放问题。我们提出了LLaDA-UI,一个基于MoE的167亿参数块级扩散视觉语言GUI智能体。LLaDA-UI采用两阶段训练流程:通用多模态预训练将原生分辨率视觉编码器与LLaDA2.0-mini-base扩散语言骨干对齐,随后在多样化的移动端、桌面端、网页和定位数据上进行GUI智能体监督微调。在广泛采用的定位基准和跨多个平台的导航基准上,LLaDA-UI大幅优于Qwen2.5-VL-7B,并在六个报告的GUI基准中的四个上超越Qwen3-VL-8B。这些结果确立了块级扩散作为多模态GUI智能体的实用生成范式。

英文摘要

Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time. However, whether dLLMs can be extended into capable multimodal GUI agents while preserving their parallel decoding advantage remains an open question. We present LLaDA-UI, a 16.7B-parameter MoE-based, block-wise diffusion vision-language GUI agent. LLaDA-UI follows a two-stage training pipeline: general multimodal pre-training aligns a native-resolution vision encoder with the LLaDA2.0-mini-base diffusion language backbone, followed by GUI-agent supervised fine-tuning on diverse mobile, desktop, web, and grounding data. Across widely adopted grounding benchmarks and navigation benchmarks spanning multiple platforms, LLaDA-UI substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks. These results establish block-wise diffusion as a practical generative paradigm for multimodal GUI agents.

↑