发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DashAct是首个细粒度诊断GUI代理在交互式仪表板分析中失败的基准,通过渐进式诊断级联评估端到端执行、动作预测和视觉定位,揭示当前模型在逐步支持下的瓶颈并提供改进指导。
AI 中文摘要
交互式仪表板要求用户通过有状态交互来揭示并连接证据。尽管图形用户界面(GUI)代理可以自动化这一过程,但现有的仪表板基准主要报告最终答案或任务成功率,对于失败是源于维持分析过程、选择操作还是定位视觉目标,它们提供的洞察有限。我们引入了DashAct,据我们所知,这是首个在同一仪表板任务中细粒度诊断这些失败的基准。DashAct包含357条经人工验证的交互轨迹,带有里程碑依赖关系和层级目标标注。其渐进式诊断级联评估端到端执行,为下一步动作预测恢复已验证的上下文,并提供目标语义和局部视图以支持视觉定位。通过逐步恢复成功所需的条件,DashAct衡量代理恢复所需的最小支持,而非对孤立技能进行评分。实验表明,即使增加支持,当前模型仍然表现挣扎。级联结果揭示了被端到端分数掩盖的瓶颈,并为改进GUI代理提供了可操作的指导。
英文摘要
Interactive dashboards require users to reveal and connect evidence across stateful interactions. Although graphical user interface (GUI) agents could automate this process, existing dashboard benchmarks primarily report final answers or task success. They provide limited insight into whether failures arise from maintaining the analytical process, selecting actions, or grounding visual targets. We introduce DashAct, to our knowledge the first benchmark to diagnose these failures at a fine-grained level within the same dashboard task. DashAct contains 357 human-verified interaction trajectories with milestone dependencies and hierarchical target annotations. Its progressive diagnostic cascade evaluates end-to-end execution, restores verified context for next-action prediction, and provides target semantics and a local view for visual grounding. By progressively restoring the conditions for success, DashAct measures the minimum support an agent needs to recover rather than scoring isolated skills. Experiments show that current models struggle even as support is added. The cascade outcomes reveal bottlenecks hidden by end-to-end scores and provide actionable guidance for improving GUI agents.