What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning
屏幕到动作中缺失了什么?面向多模态GUI推理的UI-in-the-Loop范式
机构 * Zhejiang University(浙江大学) ; ZJU-Ant Group Joint Lab of Knowledge Graph(浙大蚂蚁集团知识图谱联合实验室)
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI
AI总结 提出UI-in-the-Loop (UILoop) 范式,通过循环的屏幕-UI元素-动作过程,让多模态大模型显式学习UI元素的定位、语义和用法,实现可解释推理,并在UI理解任务上达到最优。
Comments Accepted by ACL 2026 Findings