arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27866cs.CV

Iron:用于增强通用虚拟智能体的意图对齐与回溯双学习框架

Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual Agents

Jiahe Ying, Wendong Bu, Kaihang Pan, Bingchen Miao, Siyu Chen, Wen Wang, Xueming Jiang, Juncheng Li, Siliang Tang

首次发表
浏览论文内容

中文总结 AI 辅助

Iron是用于训练GUI智能体的意图对齐与回溯双学习框架,通过逐步循环一致奖励和事后回放机制提升性能,在跨环境等任务中优于三倍数据训练的模型,未见过的网页任务获25.06%相对提升。

中文摘要 AI 辅助

实现能在各类数字环境中自动执行任务的虚拟智能体,仍是具身智能领域的关键挑战。多模态大语言模型(MLLM)虽具备更强的视觉感知与推理能力,但其智能体部署面临三大难题:数据标注成本高昂、动作-意图对齐不精确、废弃失败轨迹导致探索效率低下。为解决这些问题,我们提出Iron——一种用于训练GUI智能体的意图对齐、自改进且标注高效的框架。Iron采用新颖的双学习策略,利用逐步循环一致(SCC)奖励实现低层动作与高层意图的细粒度对齐,从而提升指令落地与意图理解能力;同时,Iron引入事后回放机制,将失败轨迹重新用于训练,提高学习效率与任务多样性。大量实验表明,经Iron训练的通用智能体在跨环境、跨设备任务中性能持续提升,表现优于用三倍数据训练的模型;Iron在未见过的网页任务上实现了25.06%的相对提升,在固有复杂任务上也观察到进一步增益,证明了构建更强大虚拟智能体的可行性。

英文摘要

Achieving virtual agents capable of automating tasks across diverse digital environments remains a pivotal challenge in Embodied AI. While Multimodal Large Language Models (MLLMs) offer enhanced visual perception and reasoning, their agentic deployment faces three challenges: costly data annotation, imprecise action-intent alignment, and inefficient exploration from discarded failed trajectories. To address these, we introduce Iron, an intent-aligned, self-improved, and annotation-efficient framework for training GUI agents. Iron employs a novel dual learning strategy that utilizes a stepwise cycle-consistent (SCC) reward to achieve fine-grained alignment between low-level actions and high-level intents, thereby improving instruction grounding and intent understanding. Concurrently, Iron introduces a hindsight reproduction mechanism to repurpose failed trajectories for training, improving both learning efficiency and task diversity. Extensive experiments demonstrate that Iron-trained generalist agents consistently improve performance on cross-environment and cross-device tasks, outperforming models trained with three times more data. Iron also achieves a substantial 25.06% relative improvement on unseen web tasks, with further gains observed on inherently complex tasks, demonstrating the feasibility of building more capable virtual agents.

发表机构

  • Fudan University(复旦大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑