AI 中文总结
针对现有GUI视觉定位模型部署后难适配未见界面且无法反思失败探索的问题,提出测试时自演进框架,结合MLLM反思器、反思引导的在线策略自蒸馏等方法,在六个基准上平均提升7.4%准确率,完善了GUI智能体自演进能力。
AI 中文摘要
GUI视觉定位是GUI智能体的基础能力。现有模型在部署后通常冻结参数,限制了其对未见界面的适应能力。尽管近期有方法尝试通过测试时强化学习来适配模型,但它们无法对失败的探索进行反思。为解决这一问题,我们提出了一种测试时自演进框架,使模型在部署后无需人工标注的真值即可改进,该框架构建了探索、评估、反思和内化的闭环。具体而言,智能体首先通过预测给定指令的定位坐标来探索未见界面;为评估这些探索,我们引入了基于多模态大语言模型(MLLM)的反思器,以评估生成的结果并提供相应的推理反思;为将反思知识内化到模型权重中,我们提出了反思引导的在线策略自蒸馏(Reflection-Guided On-Policy Self-Distillation),其通过条件自教师将高层推理转化为密集令牌级监督;此外,我们设计了对比校准方法,以防止失败探索期间不正确的自回归前缀破坏监督信号。在六个基准上的大量实验证明了我们框架的有效性,相较于基础模型实现了7.4%的平均准确率提升。据我们所知,这是首个成功利用在线策略自蒸馏进行GUI视觉定位测试时适配的工作,通过填补部署后适配的空白,我们的框架完善了GUI智能体的自演进能力,代码将被公开。
英文摘要
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon failed exploration. To overcome this, we propose a Test-Time Self-Evolving framework that enables models to improve after deployment without human-annotated ground truth. It constructs a closed-loop of Exploration, Evaluation, Reflection, and Internalization. Specifically, the agent first explores unseen interfaces by predicting grounding coordinates for given instructions. To evaluate these explorations, we introduce an MLLM-based Reflector to assess the generated results and provide the corresponding reasoning reflections. To internalize reflection knowledge into the model weights, we propose Reflection-Guided On-Policy Self-Distillation, which translates high-level reasoning into dense token-level supervision via a conditioned self-teacher. Furthermore, we design a Contrastive Calibration method to prevent incorrect auto-regressive prefixes from corrupting the supervisory signals during failed explorations. Extensive experiments across six benchmarks demonstrate our framework's effectiveness, achieving an average accuracy improvement of 7.4% over the base model. To the best of our knowledge, this is the first work to successfully exploit on-policy self-distillation for test-time adaptation in GUI visual grounding. By filling the gap in post-deployment adaptation, our framework completes the self-evolving capability of GUI agents. The code will be released.