Benchmarking LLM Judges for Mobile Agent Evaluation
面向移动智能体评估的LLM评判基准测试
Ziqiang Wang, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang
机构
*
Mila – Québec AI Institute(米拉-魁北克人工智能研究所)
;
Concordia University(康考迪亚大学)
;
University of Toronto(多伦多大学)
;
Shanghai University(上海大学)
;
McMaster University(麦克马斯特大学)
Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support
机构
*
University of Southern California(南加利福尼亚大学)
;
University of Chicago(芝加哥大学)
;
University of California, Berkeley(加利福尼亚大学伯克利分校)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Stanford University(斯坦福大学)
;
University of California, Davis(加利福尼亚大学戴维斯分校)
;
Pennsylvania State University(宾夕法尼亚州立大学)
;
Harvard University(哈佛大学)
;
University of Oxford(牛津大学)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
University of Adelaide(阿德莱德大学)
;
Tsinghua University(清华大学)
;
National University of Singapore(新加坡国立大学)
;
Peking University(北京大学)
Comments42 pages, 18 figures. Extended version of a paper presented at ICAART 2026; submitted for consideration in the ICAART 2026 post-publication selected-papers volume in Lecture Notes in Artificial Intelligence
Comments7 pages. To be published in the proceedings of 41st International Conference on Automated Software Engineering (ASE '26), October 12-16, 2026, Munich, Germany (Industry Showcase Track)