arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Vinci2:在连续自我中心视频中提供主动协助

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng, Xiang Ruan, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang

arXiv 2607.11523首次发表:更新:

发表机构

Dalian University of Technology; Alaya Lab; The University of Tokyo(大连理工大学; 阿莱雅实验室; 东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究连续自我中心视频中智能助手的主动协助问题,提出Vinci2系统,包含EgoServe基准测试和EgoMemo智能体,通过依赖上下文决策及相关技术,在新基准测试上建立强大基线且在现有测试中具竞争力。

AI 中文摘要

智能助手何时应主动发声?连续自我中心视频提供了丰富且不断演变的背景,能实现一种新的主动协助形式。现有方法要么被动等待用户查询,要么对每个检测到的事件都做出响应,未考虑用户历史、当前活动等。我们将主动协助重新构建为一个依赖上下文的决策问题。为此,我们提出了Vinci2系统,它推动设备上的助手从被动响应向主动协助发展。我们还介绍了EgoServe这个首个用于连续自我中心视频中主动协助的大规模基准测试,以及EgoMemo这种无需训练、内存增强的智能体。实验表明EgoMemo在EgoServe上建立了强大基线,在现有自我中心基准测试中也具有竞争力。

英文摘要

When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at \href{https://sitonggong.github.io/EgoServe-page/}{Vinci2}.

CommentsAccepted by ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑