arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10297cs.CV

TRACE:面向高效GUI智能体的轨迹鲁棒准入与证据排序

TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

  • Dalian University of Technology(大连理工大学)
  • OPPO Research Institute(OPPO研究院)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuhao Wang, Mu Qiao, Xindong Zhang, Yunzhi Zhuge, Lei Zhang, Huchuan Lu

AI总结:

针对GUI智能体轨迹中高分辨率截图导致的推理延迟和内存问题,提出免训练框架TRACE,结合布局先验、指令相关性和特征新颖性进行证据排序,并预留预算修复空间覆盖,通过单调KV压缩实现高效视觉令牌管理,在六个基准上验证了有效性。

AI中文摘要:

GUI智能体在轨迹展开过程中会累积高分辨率截图,导致推理延迟和内存使用增加。免训练的视觉令牌剪枝可以降低这一成本,但缓存复用引入了一个基本约束:一旦令牌被丢弃,相应的视觉证据在不重新编码的情况下无法恢复。因此,剪枝成为一个不可逆的准入决策,该决策必须在预算紧张的情况下,对未知的未来目标保持有用,同时保留可操作区域的覆盖。为应对这些挑战,我们提出了TRACE,一个免训练框架,用于轨迹鲁棒准入与覆盖感知的证据排序。具体而言,我们将与查询无关的布局派生交互先验与指令相关性和特征新颖性相结合,根据潜在未来效用和多样性对视觉证据进行排序。然后,我们预留部分预算用于分布在整个屏幕上的原生视觉令牌,在不破坏排序的情况下修复缺失的空间覆盖。这些机制共同产生一个嵌套的令牌顺序,使得保留的视觉证据能够在不同预算下单调收缩,同时在整个轨迹中保持可复用。最后,我们的单调KV压缩逐步将退役帧压缩为紧凑的会话状态,避免重复的视觉编码或剪枝。在六个GUI基准和多种模型上的大量实验验证了所提方法在紧张预算下的有效性。源代码将发布。

英文摘要:

GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usage. Training-free visual token pruning can reduce this cost, but cache reuse introduces a fundamental constraint. Once tokens are discarded, the corresponding visual evidence cannot be recovered without re-encoding. Pruning therefore becomes an \textit{irreversible admission decision} that must remain useful for unknown future targets while preserving coverage of operable regions under tight budgets. To address these challenges, we propose \textbf{\method{}}, a training-free framework for \emph{\textbf{T}rajectory-\textbf{r}obust \textbf{A}dmission and \textbf{C}overage-aware \textbf{E}vidence ordering}. Specifically, we combine a query-independent layout-derived interaction prior with instruction relevance and feature novelty to rank visual evidence according to both potential future utility and diversity. Then, we reserve part of the budget for native visual tokens distributed across the screen, repairing missing spatial coverage without breaking the ordering. Together, these mechanisms produce a nested token order, allowing retained visual evidence to shrink monotonically across budgets while remaining reusable throughout the trajectory. Finally, our monotone KV contraction incrementally contracts retired frames into compact session state, avoiding repeated visual encoding or pruning. Extensive experiments across six GUI benchmarks and diverse models verify the effectiveness of our proposed \method{} under tight budgets. The source code will be released.

↑