arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21830cs.AI

超越成功与失败:面向GUI智能体的长度感知对比学习

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

Chengyang Gu, Le Zhang, Jingbo Zhou, Yize Chen, Yu Shi, Siqi Bao, Zheng-Fan Wu, Hua Wu, Hui Xiong

首次发表
浏览论文内容

中文总结 AI 辅助

针对GUI智能体现有对比RLVR方法无法捕获轨迹细粒度质量差异的问题,提出LACL-GUI框架,引入轨迹级质量信号,在基准测试中实现性能提升。

中文摘要 AI 辅助

由多模态大语言模型(MLLMs)驱动的图形用户界面(GUI)智能体,在跨数字环境自动化任务方面展现出强大潜力,其中强化学习(RL)已成为主流训练范式。然而,Group Relative Policy Optimization(GRPO)等广泛使用的方法存在奖励梯度错位问题,导致优化效率低下且不稳定。近期研究通过将带可验证奖励的强化学习(RLVR)重新表述为对比或分类目标,解决了该问题,通过消除不良梯度行为提升了稳定性。尽管取得了这些进展,现有的对比RLVR方法主要依赖结果级监督,无法捕获同一结果类别内轨迹质量的细粒度差异。本文提出面向GUI智能体的长度感知对比学习(LACL-GUI),这是一种将轨迹级质量信号融入策略优化的对比RLVR框架。LACL-GUI在成功与失败轨迹中引入结构化偏好,鼓励简洁的成功执行,并根据与成功轨迹的偏差区分失败质量,同时保持优化稳定性。在GUI智能体基准上的实验表明,LACL-GUI提供了更有效的学习信号,且相比现有方法持续提升智能体性能,凸显了轨迹级监督在对比RLVR中的价值。

英文摘要

Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) have shown strong potential for automating tasks across diverse digital environments, where reinforcement learning (RL) has become a dominant training paradigm. However, widely used methods such as Group Relative Policy Optimization (GRPO) suffer from reward-gradient misalignment, leading to inefficient and unstable optimization. Recent work addresses this issue by reformulating RL with verifiable rewards (RLVR) as contrastive or classification-based objectives, which improve stability by eliminating problematic gradient behaviors. Despite this progress, existing contrastive RLVR methods rely primarily on outcome-level supervision and fail to capture fine-grained differences in trajectory quality within the same outcome category. In this paper, we propose Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive RLVR framework that incorporates trajectory-level quality signals into policy optimization. LACL-GUI introduces structured preferences within both successful and failed trajectories, encouraging concise successful executions and differentiating failure quality based on divergence from successful trajectories, while preserving optimization stability. Experiments on GUI agent benchmarks show that LACL-GUI provides more effective learning signals and consistently improves agent performance over prior methods, highlighting the value of trajectory-level supervision in contrastive RLVR.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Baidu Inc.(百度公司)
  • University of Alberta(阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑