面向移动智能体评估的LLM评判基准测试
Benchmarking LLM Judges for Mobile Agent Evaluation
浏览论文内容
中文总结 AI 辅助
该研究推出MobileJudgeBench基准,评估6种LLM评判器方法在移动智能体轨迹上的可靠性,发现简单基线评判器具竞争力、基准质量指标可预测评判器效用,且不同LLM后端故障特征相反。
中文摘要 AI 辅助
移动智能体基准测试越来越依赖基于大语言模型(LLM)的评判器来评估任务完成情况,但这些评判器在移动智能体轨迹上的可靠性在很大程度上未被检验。我们推出MobileJudgeBench,这是一个用于系统评估移动智能体轨迹上LLM作为评判器方法的基准。我们的基准包含931条人工标注的轨迹,覆盖6个移动智能体基准、4个智能体模型和68个应用程序。我们在多个LLM后端上评估了6种评判器方法(5种改编自SPA-Bench、A3的两种模式、AndroidArena和AgentRewardBench,加上我们设计的一个简单基线)。我们的实验揭示了三个关键发现:第一,带有采样截图的简单基线评判器与专用方法具有竞争力,且常常超过它们,这表明更复杂的评判器流程并不总能提升评判质量;在具有竞争力的方法中,LLM主干是主要驱动因素。第二,基准质量指标可可靠预测现实世界的评判器效用:它们既与评估的智能体排序保真度相关,也与评判器作为在线策略强化学习的奖励信号时的下游性能相关。第三,对两种LLM后端的故障分析揭示了性质相反的故障特征,一种保守,另一种宽松,这与主干的精确率-召回率特性相关。
英文摘要
Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge methods on mobile agent trajectories. Our benchmark comprises 931 human-annotated trajectories spanning 6 mobile agent benchmarks, 4 agent models, and 68 apps. We evaluate 6 judge methods (five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline we design) across multiple LLM backends. Our experiments reveal three key findings. First, a simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods, indicating that more elaborate judge pipelines do not consistently improve judge quality; among competitive methods, the LLM backbone is the primary driver. Second, benchmark quality metrics reliably predict real-world judge utility: they correlate with both agent ranking fidelity for evaluation and downstream performance when judges serve as reward signals for on-policy reinforcement learning. Third, failure analysis across two LLM backends uncovers qualitatively opposite failure profiles, one conservative and the other permissive, linked to the backbone's precision-recall characteristics.
发表机构
- Mila – Québec AI Institute(米拉-魁北克人工智能研究所)
- Concordia University(康考迪亚大学)
- University of Toronto(多伦多大学)
- Shanghai University(上海大学)
- McMaster University(麦克马斯特大学)
机构由 AI 辅助整理,请以论文原文为准。