相关性并不意味着适用性:面向个人GUI代理的经验激活
Relevance Does Not Imply Applicability: Experience Activation for Personal GUI Agents
浏览论文内容
中文总结 AI 辅助
本文发现检索相关历史对个人GUI代理的决策帮助主要集中在前几步,提出无需训练的ExpActivator框架,通过匹配当前屏幕与历史状态激活适用经验,将轨迹内步骤成功率平均提高28%。
中文摘要 AI 辅助
个人图形用户界面(GUI)代理依赖交互历史来从模糊指令中推断用户意图,并预判重复出现的例行程序。现有方法检索与任务相关的历史并将其附加到策略的上下文中,隐含地假设与任务相关的经验对其中的每个决策仍然有用。我们发现这种帮助在很大程度上消耗在第一个决策上:检索到的历史显著改善了情节的起始步骤,但在剩余的90%步骤中几乎没有持续益处,并且在是否应提出主动建议方面提供的指导很弱。相关记录可能告诉代理从哪里开始,但无法告知过去的哪个动作适用于当前屏幕,或某个例行程序现在是否该执行。根本问题在于相关性并不意味着适用性:相关性在任务层面确定,而适用性取决于决策时的情境。因此,我们将个性化重新定义为经验激活,并引入ExpActivator,这是一个无需训练的框架,仅激活适用于当前情境的经验。在执行过程中,ExpActivator将每个新屏幕与冻结的GUI骨干网络潜在空间中的历史状态进行匹配,并提供相应的动作作为参考。在执行前,仅当当前时间和场景提供足够支持时,它才激活重复出现的意图,否则弃权(不执行)。在四个GUI骨干网络上,ExpActivator将轨迹内步骤成功率平均提高28%,在每个骨干网络上实现了最佳个性化执行,同时使用的历史令牌约为五分之一,并达到最强主动基线马修斯相关系数的约2.3倍。经验在激活之处发挥作用,而非附加之处。
英文摘要
Personal Graphical User Interface (GUI) agents rely on interaction history to infer what a user wants from ambiguous instructions and to anticipate recurring routines. Existing approaches retrieve task-relevant history and append it to the policy's context, implicitly assuming that experience relevant to a task remains useful for each decision within it. We find that this help is largely spent at the first decision: retrieved history strongly improves the opening step of an episode, yet provides little sustained benefit over the remaining 90\% of steps, and offers weak guidance on whether a proactive suggestion is warranted. A relevant record may tell the agent where to begin, but not which past action applies to the current screen or whether a routine is due now. The underlying issue is that relevance does not imply applicability}: relevance is determined at the task level, whereas applicability depends on the situation at decision time. We therefore recast personalization as experience activation and introduce ExpActivator, a training-free framework that activates only the experience applicable to the current situation. During execution, ExpActivator matches each new screen to historical states in the frozen GUI backbone's latent space and supplies the corresponding action as a reference. Before execution, it activates a recurring intent only when the current time and scenario provide sufficient support, and otherwise abstains. Across four GUI backbones, ExpActivator improves within-trajectory step success by 28\% on average, achieves the best personalized execution on every backbone while using about one-fifth as many history tokens, and reaches approximately 2.3$\times$ the Matthews correlation coefficient of the strongest proactive baseline. Experience pays where it is activated, not where it is appended.