发表机构
ByteDance Seed; M-A-P(字节跳动Seed; M-A-P)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
xDailyBench基准包含248个真实生活任务,评估11个模型在显性与隐性需求上的表现,发现隐性推断是主要瓶颈,最佳模型得分75.6%。
AI 中文摘要
大语言模型(LLMs)越来越多地用于日常辅助,然而现有基准仅部分反映了用户在实践中的自然请求。现实世界的请求通常是开放式的、随意指定的且依赖上下文的,要求模型不仅遵循明确的指令,还要从用户背景和情境中推断未言明的需求。我们引入了xDailyBench,一个包含248个精心策划任务的基准,涵盖个人生活、白领工作、学习与研究以及跨领域活动等51个场景。这些任务基于用户实际完成或真正打算借助AI完成的请求,并通过涵盖显性和隐性要求的细粒度二元评分标准进行评估。我们在标准化智能体设置下评估了11个前沿模型。最佳模型在任务级得分达到75.6%,而所有模型在隐性要求上的表现显著差于显性要求,差距不少于9个百分点。这些结果表明,隐性要求推断是可靠满足现实世界日常用户需求的持续瓶颈。
英文摘要
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.