arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06721cs.CVcs.AI

自我中心视觉中的伴侣式问答辅助

Companion-style QA Assistance in Ego-Vision

  • National University of Singapore(新加坡国立大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Hangyu Qin, Junbin Xiao, Shenglang Zhang, Angela Yao

AI总结:

提出BuddyVQA基准和MyBuddy助手,通过多模态思维链推理解决自我中心视频中的伴侣式问答,显著提升基础模型性能。

AI中文摘要:

AI伴侣被设想为始终在线的助手,支持用户的日常生活。为此,我们引入了BuddyVQA,一个用于自我中心流式视频的伴侣式问答(QA)基准。BuddyVQA包含21.6K个问题,关联到1,012个长时自我中心视频中的6K个高光时刻。它具备两个在日常第一人称问答辅助中常见但在现有VideoQA基准中大多被忽视的关键特征:自我中心指示表达和交互式链式问题(例如,“它在哪里?”,“如何到达那里?”)。这些要求模型通过解析视觉代词,在自我中心视觉和问答内容的上下文中推断用户的现场意图,两者都基于长时流式设置。为应对这些挑战,我们提出了MyBuddy,一个伴侣式问答助手,其突出特点是多模态思维链推理机制,基于历史问答和视觉内容推断最终答案。此外,设计了额外的问题过滤器和多级记忆,以促进流式问答设置下的高效问答和视觉信息检索。实验表明,MyBuddy显著提升了基础模型在BuddyVQA上的性能。此外,这些增益泛化到其他流式和常见视频问答基准,证明了我们方法的适用性和有效性。我们的代码和数据集可在以下网址获取:此https URL。

英文摘要:

AI companions are envisioned as always-on assistants that support users in daily life. With this regard, we introduce BuddyVQA, a benchmark for companion-style question answering (QA) on egocentric streaming video. BuddyVQA contains 21.6K questions linked to 6K highlight moments across 1,012 long, egocentric videos. It features two key characteristics that are common in daily first-person QA assistance but are largely overlooked in existing VideoQA benchmarks: ego-deictic expressions and interactively chained questions (e.g., "Where is it?", "How to get there?"). These require models to infer a user's in-situation intent by resolving visual pronouns in the context of egocentric visual and QA contents, with both grounded in a long-form streaming setting. To tackle the challenges, we propose MyBuddy, a companion-style QA assistant that highlights a multimodal chain-of-thought reasoning mechanism to infer the final answer based on the historical QA and visual content. An additional question filter and multi-level memory are designed to facilitate efficient QA and visual information retrieval under streaming QA settings. Experiments show that MyBuddy significantly enhances the performance of foundation models on BuddyVQA. Moreover, these gains generalize to other streaming and common video QA benchmarks, demonstrating the applicability and effectiveness of our approach. Our code and dataset are available at https://github.com/QHUni/BuddyVQA

补充信息

↑