CommentsThis manuscript was inadvertently made publicly available before all necessary internal review processes had been completed. The authors are withdrawing the manuscript
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
FieldWorkArena:面向真实作业任务的代理AI基准测试
Jun Takahashi, Atsunori Moteki, Akiyoshi Uchida, Shoichi Masui, Fan Yang, Kanji Uchino, Yueqi Song, Yonatan Bisk, Graham Neubig, Ikuo Kusajima, Yasuto Watanabe, Hiroyuki Ishida, Koki Nakagawa, Shan Jiang
机构
*
Fujitsu Limited(富士通株式会社)
;
Fujitsu Research of America(富士通美国研究部)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Master’s Student, The University of Tokyo(东京大学硕士研究生)
;
Agent Research Collective(代理研究集体)
机构
*
York University(约克大学)
;
Bangladesh University of Business and Technology(孟加拉国商业与技术大学)
;
RBC, Canada(加拿大RBC)
;
Nanyang Technological University(南洋理工大学)
;
Salesforce AI Research(Salesforce人工智能研究)
Comments9 pages. v2: results updated to July 2026 leaderboard (17 models). Accepted at the 2nd Workshop on Knowledge-Intensive Multimodal Reasoning (KnowledgeMR) at CVPR 2026 (non-archival), under the former title "PDFParse: A Benchmark for Grounded Multimodal Reasoning over Professional PDF Documents". Dataset: https://huggingface.co/datasets/surgeai/GDP.pdf ; Code: https://github.com/surge-ai/gdp-pdf
机构
*
School of Information Science and Technology, University of Science and Technology of China(科学技术大学信息科学与技术学院)
;
ByteDance Intelligent Creation(字节跳动智能创作)
;
School of Computer Science and Technology, Harbin Institute of Technology (Weihai)(哈尔滨工业大学(威海)计算机科学与技术学院)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院)