arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11434cs.AIcs.CLcs.CV

面向移动智能体评估的LLM评判基准测试

Benchmarking LLM Judges for Mobile Agent Evaluation

Ziqiang Wan, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出MobileJudgeBench基准,评估6种LLM评判器方法在移动智能体轨迹上的可靠性,发现简单基线评判器具竞争力、基准质量指标可预测评判器效用,且不同LLM后端故障特征相反。

中文摘要 AI 辅助

移动智能体基准测试越来越依赖基于大语言模型(LLM)的评判器来评估任务完成情况,但这些评判器在移动智能体轨迹上的可靠性在很大程度上未被检验。我们推出MobileJudgeBench,这是一个用于系统评估移动智能体轨迹上LLM作为评判器方法的基准。我们的基准包含931条人工标注的轨迹,覆盖6个移动智能体基准、4个智能体模型和68个应用程序。我们在多个LLM后端上评估了6种评判器方法(5种改编自SPA-Bench、A3的两种模式、AndroidArena和AgentRewardBench,加上我们设计的一个简单基线)。我们的实验揭示了三个关键发现:第一,带有采样截图的简单基线评判器与专用方法具有竞争力,且常常超过它们,这表明更复杂的评判器流程并不总能提升评判质量;在具有竞争力的方法中,LLM主干是主要驱动因素。第二,基准质量指标可可靠预测现实世界的评判器效用:它们既与评估的智能体排序保真度相关,也与评判器作为在线策略强化学习的奖励信号时的下游性能相关。第三,对两种LLM后端的故障分析揭示了性质相反的故障特征,一种保守,另一种宽松,这与主干的精确率-召回率特性相关。

英文摘要

Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge methods on mobile agent trajectories. Our benchmark comprises 931 human-annotated trajectories spanning 6 mobile agent benchmarks, 4 agent models, and 68 apps. We evaluate 6 judge methods (five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline we design) across multiple LLM backends. Our experiments reveal three key findings. First, a simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods, indicating that more elaborate judge pipelines do not consistently improve judge quality; among competitive methods, the LLM backbone is the primary driver. Second, benchmark quality metrics reliably predict real-world judge utility: they correlate with both agent ranking fidelity for evaluation and downstream performance when judges serve as reward signals for on-policy reinforcement learning. Third, failure analysis across two LLM backends uncovers qualitatively opposite failure profiles, one conservative and the other permissive, linked to the backbone's precision-recall characteristics.

发表机构

  • Mila – Québec AI Institute(米拉-魁北克人工智能研究所)
  • Concordia University(康考迪亚大学)
  • University of Toronto(多伦多大学)
  • Shanghai University(上海大学)
  • McMaster University(麦克马斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑