arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有经验都属于权重:面向自改进GUI代理的组件路由

Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents

Beining Wu, Zihao Ding, Jun Huang

arXiv 2610.01787首次发表:更新:

发表机构

South Dakota State University(南达科他州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出组件路由方法,将GUI代理经验拆分为不同组件并分别路由到上下文或权重,通过规则预测最优目的地,显著提升自改进代理性能。

AI 中文摘要

自改进的GUI代理保留其产生的轨迹,并通过微调或检索到提示中将其返回给代理,而比较这两种目的地的研究结果并不一致。我们将此归因于经验的单位:一条轨迹将具有不同属性的项目捆绑在一起,因此关于捆绑包的结论取决于其构成。为解决这一问题,(i) 我们引入了组件路由,将经验拆分为定位器、程序、状态事实和教训,并将每个组件发送到上下文或权重中,在三个骨干家族、两个环境和三个种子上对相同的项目进行比较。一个池有两个目的地:定位器和教训在权重中获胜,程序和状态事实在上下文中的表现更优。(ii) 我们在任何训练之前,根据两个属性(重复性和状态条件性)拟合了一个规则;该规则在24个单元格中的24个中恢复了留出骨干家族的目的地,两次干预将组件推向边界,按规则路由击败了所有整条轨迹的基线,并且平均比每个骨干的更好单一目的地高出+3.5分。(iii) 我们确定了训练和生产者-消费者差异如何改变两个目的地的价值:在将相同组件写入权重后,注释读出减少,对于重复最多的项目减少最多,上下文增益随着信息差距的增加而增加,权重增益随着策略差距的增加而减少。代码和数据将发布。

英文摘要

Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a conclusion about the bundle depends on its mix. To address this, (i) we introduce component routing, which splits the experience into locators, procedures, state facts and lessons and sends each component to the context or to the weights, compared on the same items across three backbone families, two environments and three seeds. One pool has two destinations: locators and lessons win in the weights, procedures and state facts in the context. (ii) We fit a rule in two properties measured before any training, recurrence and state-conditionality; it recovers the destination of a held-out backbone family in 24 of 24 cells, two interventions move a component toward the boundary, and routing by the rule beats every whole-trajectory baseline and, by +3.5 points on average, the better single destination of each backbone. (iii) We identify how training and producer-consumer differences change the value of the two destinations: note readout decreases after the same component is written into the weights, most for the items that recur most, context gains increase with the information gap, and weights gains decrease with the policy gap. Code and data will be released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑