arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

xDailyBench:面向现实生活问题的专业咨询大语言模型基准测试

xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems

Yongchang Peng, Qingshui Gu, Liya Zhu, Ge Zhang, Duo Wang, Haodong Wang, Jingzhe Ding, Tianhao Yu, Letian Gao, Yongjie Zhong, Chaoxin Li, Zixin Su, Jinchao Tao, Xingyu Ma, Xin'ao Guo, Feng Tian, Shiyuan Dong, Xiaoyan He, Sen Liu, Xin Chen, Jiajun Li, Zejia Zhang, Xi Lin, Wen Zhang, Yi Zhu, Duju Zeng, Xiang Gao, Yunyang Wang, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang

arXiv 2609.07784首次发表:更新:

发表机构

ByteDance Seed; M-A-P(字节跳动Seed; M-A-P)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

xDailyBench基准包含248个真实生活任务,评估11个模型在显性与隐性需求上的表现,发现隐性推断是主要瓶颈,最佳模型得分75.6%。

AI 中文摘要

大语言模型(LLMs)越来越多地用于日常辅助,然而现有基准仅部分反映了用户在实践中的自然请求。现实世界的请求通常是开放式的、随意指定的且依赖上下文的,要求模型不仅遵循明确的指令,还要从用户背景和情境中推断未言明的需求。我们引入了xDailyBench,一个包含248个精心策划任务的基准,涵盖个人生活、白领工作、学习与研究以及跨领域活动等51个场景。这些任务基于用户实际完成或真正打算借助AI完成的请求,并通过涵盖显性和隐性要求的细粒度二元评分标准进行评估。我们在标准化智能体设置下评估了11个前沿模型。最佳模型在任务级得分达到75.6%,而所有模型在隐性要求上的表现显著差于显性要求,差距不少于9个百分点。这些结果表明,隐性要求推断是可靠满足现实世界日常用户需求的持续瓶颈。

英文摘要

Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑