arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VibeLifeBench:你的生命智能体能在真实生活环境中保持主动与持续性吗?

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Xiaohongshu Dots Studio, Evolvent AI

arXiv 2608.10875首次发表:更新:

发表机构

Xiaohongshu Dots Studio; Evolvent AI(小红书Dots工作室; Evolvent AI公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出VibeLifeBench基准,评估7个前沿LLM智能体在200项长周期日常生活任务中的表现,发现其主动与持续性不足,将开源相关任务、环境与评估框架。

AI 中文摘要

大型语言模型(LLM)智能体越来越多地被部署为个人助手。然而,现有的评估大多使用静态环境中简短、自包含的请求,而日常生活辅助则不同:一项任务会持续数周而非数分钟,在智能体未被提示时世界仍在不断变化,许多约束从未明确说明。仅对眼前请求做出回应的智能体将无法完成这类任务,相反需要的是主动且一致的智能体:它能自行决定何时行动、何时询问、何时弃权(不执行),能注意到无人宣布的变化,能从第一天到最后一天保持计划连贯。目前没有任何基准能衡量这一点。我们推出VibeLifeBench,这是一个涵盖10个日常生活领域的200项长周期任务的基准,每项任务都是在包含22个模拟服务的虚拟世界中设定的多周时间线;世界按自身时钟推进,其许多变化是无声的,只有重新审视世界的智能体才能发现这些变化。每项任务通过细粒度的加权检查进行评分,仅读取智能体实际留下的内容,涵盖最终状态、行动的及时性以及是否遵守了隐含约束。我们评估了7个前沿模型,所有模型得分都很低,这表明当前智能体距离协助真实生活还有很大差距。我们将开源所有任务、环境和评估框架。

英文摘要

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑