arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FocusMem:对潜在GUI记忆中的内容、读出和信任进行因子分解

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

Zhuoran Zhang, Bowen Li, Jingcheng Ju, Yang Shi, Qixun Wang, Haotian Wang, Wei Chen, Tengjiao Wang

arXiv 2608.04530首次发表:更新:

AI 中文总结

FocusMem将潜在GUI记忆的内容、读出和信任因子分解,解决现有固定记忆方法的细节丢失、阶段适配差和误导问题,在五个GUI智能体基准测试中性能优于基线与现有方法。

AI 中文摘要

智能体(GUI agents)必须记住早期任务中的有用经验以及当前交互中未完成的进度。潜在记忆(Latent memory)提供了一种紧凑的解决方案,它将多模态轨迹压缩为少量连续令牌。然而,现有方法通常将每个轨迹映射到一个固定的记忆块,并且主要通过下一个动作监督进行训练,这造成了三个实际问题:压缩过程中可能丢失重要细节;同一个记忆块必须服务于不同的决策阶段;检索到的不相关轨迹仍可能误导智能体。我们提出FocusMem,它在紧凑的潜在记忆接口内分离了这些职责:角色感知的内容基础(role-aware content basis)促使情景记忆保留可复用经验,工作记忆保留任务进度;状态条件读出(state-conditioned readout)生成存储证据的决策特定视图;轻量信任门(lightweight trust gate)可抑制与当前步骤无关的记忆块。所有组件在GUI策略保持冻结的情况下进行训练。在五个GUI智能体基准测试中,FocusMem始终优于完全匹配的仅动作固定记忆基线以及现有的潜在记忆适配方法。进一步分析显示,语义和功能监督保留了互补信息;随着周围轨迹上下文的增长,状态条件读出表现出更强的鲁棒性;信任门减少了注入的不相关情景证据造成的损害。这些结果表明,有效的潜在记忆不仅取决于对过去交互的压缩,还取决于保留什么、暴露什么以及允许什么。

英文摘要

GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact solution by compressing multimodal trajectories into a few continuous tokens. Existing methods, however, usually map each trajectory to one fixed memory block and train it mainly through next-action supervision. This creates three practical problems: important details may be lost during compression, the same memory block must serve different decision stages, and irrelevant retrieved trajectories may still mislead the agent. We introduce FocusMem, which separates these responsibilities within a compact latent-memory interface. A role-aware content basis encourages episodic memory to retain reusable experience and working memory to retain task progress. A state-conditioned readout generates a decision-specific view of the same stored evidence, while a lightweight trust gate can suppress memory blocks that appear irrelevant to the current step. All components are trained while the GUI policy remains frozen. Across five GUI-agent benchmarks, FocusMem consistently outperforms a fully matched action-only fixed-memory baseline and prior latent memory adaptations. Further analysis shows that semantic and functional supervision preserve complementary information, state-conditioned readout is more robust as surrounding trajectory context grows, and the trust gate reduces the harm caused by injected irrelevant episodic evidence. These results show that effective latent memory depends not only on compressing past interaction, but also on what is retained, what is exposed, and what is allowed.

Comments36 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑