arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EMBER-Bench:长时程具身任务中跨事件因果记忆的基准测试

EMBER-Bench: Benchmarking Cross-Event Causal Memory in Long-Horizon Embodied Tasks

Aoyang Cai, Boning Zhao, Shaoxuan Xie, Dahui Gao, Huan Yang, Zhongyuan Wang, Zhiwei Yu, Guocai Yao

arXiv 2610.05013首次发表:更新:

发表机构

Tsinghua University; The University of Hong Kong; Beijing Academy of Artificial Intelligence(清华大学; 香港大学; 北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有具身记忆基准忽视跨事件因果推理的问题,提出EMBER-Bench基准,含189个任务和699个问答对,评估发现最高准确率61.2%,远低于人类98.3%,揭示智能体从历史提取因果并约束当前行动仍是关键难点。

AI 中文摘要

终身物理智能体必须在长时间交互中进行推理,在这些交互中,过去的事件在从视野中消失很久之后仍持续影响世界。除了回忆发生了什么,智能体还必须推断历史如何改变当前状态并约束未来行动。然而,现有的具身和视频记忆基准主要关注历史检索和总结,使得这种依赖历史的因果推理未得到充分探索。我们引入了EMBER-Bench,一个用于长时程具身任务中跨事件因果推理的自我中心基准,为此我们新创建了任务设计、视频录制和数据标注。它包含189个家务任务和699个问答对,涵盖任务进度、失败恢复、外部干预以及具有远距离依赖和先决条件的复合长时程任务,并带有细粒度的事件和因果链标注。EMBER-Bench评估两个方向的推理:下一步行动预测从历史中选择下一个行动,以及因果回溯,在给定该行动的情况下,识别使其成为必要的历史事件。输入消融实验向视频中添加行动日志或特权因果后果标注,以指示模型未能使用哪种历史信息。在评估的16个模型中,最高总体准确率为61.2%,而两位人类评估者的平均准确率为98.3%。在成对的决策点上,正确的回溯与正确的下一步行动预测不相关。添加行动日志带来1.6个百分点的提升,而在此基础上添加因果后果标注额外带来13.0个百分点的提升。这些结果表明,从过去事件中提取因果信息并将其转化为对当前行动的约束,仍然是长时程具身智能体的关键难点。项目页面:此https URL

英文摘要

Lifelong physical agents must reason over extended interactions where past events continue to shape the world long after they disappear from view. Beyond recalling what happened, agents must infer how history changes the current state and constrains future actions. Yet existing embodied and video-memory benchmarks largely focus on historical retrieval and summary, leaving such history-dependent causal reasoning underexplored. We introduce EMBER-Bench, an egocentric benchmark for cross-event causal reasoning in long-horizon embodied tasks, for which we newly created the task design, video recording, and data annotation. It contains 189 household tasks and 699 QA pairs, spanning task progress, failure recovery, external interventions, and compound long-horizon tasks with distant dependencies and prerequisites, with fine-grained event and causal-chain annotations. EMBER-Bench evaluates reasoning in both directions: next-action prediction selects the next action from history, and causal traceback, given that action, identifies the historical event that makes it necessary. Input ablations that add action logs or privileged cause-and-consequence annotations to the video indicate which kind of historical information models fail to use. Among the 16 evaluated models, the highest overall accuracy is 61.2%, compared with a mean of 98.3% across two human evaluators. At paired decision points, correct traceback is not associated with correct next-action prediction. Adding action logs yields a gain of 1.6 points, whereas cause-and-consequence annotations yield an additional gain of 13.0 points on top of that. These results suggest that extracting causal information from past events and converting it into constraints on current actions remains a key difficulty for long-horizon embodied agents. Project Page: https://zhaoalexgoat.github.io/EMBER-Bench/

Comments26 pages, 4 figures, 14 tables. The first two authors contributed equally. Project page: https://zhaoalexgoat.github.io/EMBER-Bench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑