arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解耦“什么”与“在哪里”:小型GUI定位模型应如何接收动作类型?

Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?

Aadi Chauhan, Arthur Ilyasov

arXiv 2610.07444首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过微调Qwen2-VL-2B,比较五种向小型GUI定位模型传递动作类型的方法,发现辅助损失、加性嵌入和提示词优于基线,且提升主要源于避免预处理缺陷,而非空间先验。

AI 中文摘要

GUI智能体决定采取哪个动作以及在哪里执行该动作;我们研究小型定位模型应如何接收动作类型。通过在Android in the Wild数据集上使用LoRA微调Qwen2-VL-2B,我们在匹配的数据、计算量和解码条件下,将扁平基线模型与五种提供动作类型的方式进行比较:辅助损失、硬路由动作词、加性学习嵌入、前置学习标记以及将类型写入提示。使用五种随机种子、基于情节聚类的自助法和种子级配对检验,在混合流上的排名清晰:辅助损失、加性嵌入和提示词各自比基线获得5到7个hit@0.10点的提升,而硬路由和前置标记与基线无显著差异。这些提升大部分是防止我们预处理选择带来的保护,而非空间先验。我们的序列化器将AITW记录的类型事件的屏幕外触摸点钳制到原点;该类别降低了基线的点击定位性能,移除它使基线提升近7个点,之后没有任何机制的成功率超过基线,且置信区间排除了2个点的效应,尽管辅助损失仍缩短了平均未命中距离;在点击和滑动流上,没有任何方法有帮助。这种提升是否超越单一序列化方式仍待探索。对于部署,使用预测类型而非真实类型时,该流程相对基线的优势尚未确立(+0.016,95%置信区间[-0.017, +0.052]),且错误的类型会使所有在推理时条件化的模型崩溃。前置标记在共享学习率下没有帮助,其行几乎未从初始化移动;以十倍学习率训练时,它达到其他三种方法的水平,但三种随机种子未能确立其优势。我们还记录了一个静默失败:通过inputs_embeds注入条件使Qwen2-VL对图像标记回退到1维位置,损失了9个点。

英文摘要

A GUI agent decides which action to take and where to take it; we ask how a small grounding model should receive the action type. Fine-tuning Qwen2-VL-2B with LoRA on Android in the Wild, we compare a flat baseline with five ways of supplying the type under matched data, compute, and decoding: an auxiliary loss, a hard-routed action word, an additive learned embedding, a prepended learned token, and the type written into the prompt. With five seeds, an episode-clustered bootstrap, and seed-level paired tests, the ranking on a mixed stream is clear: the auxiliary loss, the additive embedding, and the prompt word each gain five to seven hit@0.10 points over the baseline, while hard routing and the prepended token are not distinguishable from it. Much of that gain is protection from a preprocessing choice of ours rather than a spatial prior. Our serializer clamps the off-screen touch point AITW records for type events to the origin; that class degrades the baseline's click grounding, and removing it lifts the baseline by nearly seven points, after which no mechanism's hit rate beats it and the intervals exclude a two-point effect, though the auxiliary loss still shortens the average miss; on a stream of taps and swipes none helps. Whether this generalizes beyond one serialization is open. For deployment, the pipeline's margin over the baseline with predicted rather than gold types is not established (+0.016, 95% interval [-0.017, +0.052]), and a wrong type collapses every model conditioned at inference. The prepended token does not help at the shared learning rate, where its rows barely move from initialization; trained ten times faster it reaches the level of the other three, with a margin three seeds do not establish. We also document a silent failure: injecting conditioning through inputs_embeds makes Qwen2-VL fall back to 1-D positions for image tokens, costing nine points.

Comments18 pages, 4 figures, 12 tables. Code and per-example logs are available at https://github.com/aadcha/action-conditioned-gui-agent

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑