发表机构
University of Michigan; ByteDance; Osaka University; Amazon; University of Pennsylvania(密歇根大学; 字节跳动; 大阪大学; 亚马逊; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SimLife平台及SimLife-BP基准,用于评估智能体从长期家庭观察中推断行为规则的能力,发现现有模型缺乏深层规则理解,揭示了长上下文模式理解是具身智能体的关键瓶颈。
AI 中文摘要
理解长期的人类行为要求智能体不仅推断人们当下的需求,还要理解习惯如何形成、为何重复以及何时改变。我们引入了SimLife,一个可扩展的平台,用于模拟具有丰富视觉观察、真实动作日志以及带音频的合成对话的长期家庭生活。基于SimLife构建的SimLife-BP评估了长上下文模式理解:即从数周或数月日常观察中推断潜在行为规则的能力。该基准包含106个片段,平均每个片段15.49小时和38.57个游戏内天数,以及1,439个问答对。每个任务在不同规则提示级别下考察直接、反事实、噪声和逆向推理。评估前沿模型和架构后,我们发现当前模型往往实现表面层次的预测而缺乏全面的规则理解,依赖频率启发式而非基于证据的if-then推理,并且在行为模式变化时难以适应。这些发现表明,长上下文模式理解仍然是未来具身智能体的主要瓶颈,而SimLife为研究日常人机交互中的记忆、个性化、适应性和长期规划开辟了更广阔的空间。
英文摘要
Understanding humans over long horizons requires agents to infer not only what people need in the moment, but also how routines form, why they repeat, and when they change. We introduce SimLife, a scalable platform for simulating long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues with audio. Built on SimLife, SimLife-BP evaluates long-context pattern understanding: the ability to infer latent behavioral rules from weeks or months of everyday observations. The benchmark contains 106 episodes averaging 15.49 hours and 38.57 in-game days, and 1,439 question-answer pairs. Each task probes direct, counterfactual, noisy, and inverse reasoning under different levels of rule hints. Evaluating frontier models and architectures, we find that current models often achieve surface-level prediction without comprehensive rule understanding, rely on frequency-based heuristics rather than if-then reasoning over evidence, and struggle to adapt when behavioral patterns change. These findings suggest that long-context pattern understanding remains a major bottleneck for future embodied agents, while SimLife opens a broader space for studying memory, personalization, adaptation, and long-horizon planning in everyday human-AI interaction.
CommentsCOLM 2026 Learning from Situated and Embodied Interaction Workshop