CausalCache:面向长视界GUI智能体的条件高保真恢复
CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents
浏览论文内容
中文总结 AI 辅助
CausalCache针对长视界GUI智能体的视觉上下文预算问题,通过重新分配预算的HGKV适配器等技术,在桌面、OSWorld-Verified及跨应用移动基准上显著提升了任务成功率。
中文摘要 AI 辅助
长视界GUI智能体可以将完整交互轨迹以文本动作记录的形式低成本保留,但仅向策略暴露少量过去事件的高保真像素信息。我们将此问题形式化为条件保真恢复:每个事件以仅摘要形式保留并链接到存档截图,而活跃视觉上下文预算$B$限制了可被提升为摘要加图像形式的事件数量。Recent-$B$将每个槽位都分配给最新事件,而CausalCache则在完整轨迹上重新分配相同的$B$次提升,仅当远距事件具有更高的条件边际效用时才驱逐最近的图像。其历史门控键/值(HGKV)适配器仅修改恢复的历史图像令牌,且在无历史图像时完全绕过。匹配预算替换组和按臂锚定的双重差分监督使均匀历史放大的价值为零;随后,预算感知选择器选择要恢复的摘要事件。在桌面环境中,冻结策略未表现出对任务相关存档截图的可靠偏好,其优先于它将替换的最近帧;HGKV在预指定的漂移包络内恰好学习到该选择性。在OSWorld-Verified上,将历史恢复至高保真度比仅摘要记忆带来约13个成功点的提升,而相同预算分配则无差异。在跨应用移动基准上零样本测试,CausalCache比相同预算的Recent分配显著提升整体成功率(全部列表提升+3.7个点),且增益集中在应有的场景:在基准元数据构建时固定的内存关键拆分上提升+8.6个点,对匹配对照组无可检测影响,且存在显著的按方法划分的交互作用。
英文摘要
Long-horizon GUI agents can retain complete action histories as compact text, but only a few historical screenshots fit in active context. We formulate this as budgeted fidelity restoration: every event remains summarized, while a fixed budget $B$ determines which events regain their archived screenshots. Recent-$B$ assigns all visual slots to the latest events. CausalCache instead scores the complete history and swaps in an older event only when its predicted utility exceeds that of a recent event. A history-gated key/value adapter modifies only restored history-image tokens and is exactly bypassed when no history image is active, preserving current-screen processing. The adapter and selector are trained with matched-budget interventions on desktop trajectories and evaluated zero-shot on mobile. On OSWorld-Verified, activating historical screenshots improves success by about $13$ percentage points over summary-only memory. Under the official $15$-step limit, CausalCache and Recent-$4$ are statistically indistinguishable; in a $30$-step diagnostic, CausalCache achieves $46.7\%$ success versus $42.4\%$ ($+4.3$ points). Zero-shot on $117$ MobileWorld tasks, CausalCache improves over Recent-$4$ from $30.2\%$ to $36.8\%$. The gain is concentrated on a pre-defined cross-app memory-candidate split ($30.6\%$ vs. $19.4\%$, $+11.2$ points), while single-app controls show no detectable difference ($43.6\%$ vs. $42.4\%$). These results show that selecting which past events regain pixels is more effective than spending a fixed visual budget entirely on recency.