发表机构
University of Tehran(德黑兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述梳理了用于第一人称视角视频理解的视觉-语言模型的发展,分析其挑战、研究方向与局限,明确了可部署具身智能的关键优先方向。
AI 中文摘要
第一人称视角视频从佩戴者的视角捕捉活动,提供了人类注意力、手-物交互以及目标导向行为的直接视图。这一视角对于可穿戴智能、辅助系统、人机交互以及具身智能愈发重要,但也带来了自我运动、遮挡、小型活动物体、视角依赖外观以及长程时间依赖等挑战。视觉-语言模型(VLMs)通过将视觉观测与语义知识、自然语言监督相连接,为应对这些挑战提供了有前景的基础。本综述对用于第一人称视角视频理解的VLMs进行了批判性回顾,梳理了从传统识别架构到多模态基础模型再到具身系统的发展历程。我们围绕任务、数据集、手-物交互理解、时间推理、帧与片段选择、多模态表示学习、提示、语义对齐以及模型适应来组织文献,特别关注基于图和以对象为中心的推理,将其作为随时间建模手、对象、动作与场景上下文之间关系的机制。我们进一步探究了第一人称感知与多模态基础模型如何支持可穿戴辅助、机器人技能学习、人机迁移以及具身决策。在所有被综述的文献中,一个一致的局限显现:当前模型识别可见对象的可靠性高于对演化中的交互、动作以及用户意图的识别,尤其是在长时长活动中。因此,我们确定了时间接地推理、感知交互的监督、高效长视频处理、多模态融合、图增强表示、跨域泛化、隐私以及可信评估作为可部署具身智能的关键优先方向。
英文摘要
Egocentric video captures activities from the wearer's perspective, providing a direct view of human attention, hand--object interaction, and goal-directed behavior. This perspective is increasingly important for wearable intelligence, assistive systems, human--robot interaction, and embodied AI, yet it introduces challenges including ego-motion, occlusion, small active objects, viewpoint-dependent appearance, and long-range temporal dependencies. Vision--language models (VLMs) offer a promising foundation for addressing these challenges by linking visual observations with semantic knowledge and natural-language supervision. This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems. We organize the literature around tasks, datasets, hand--object interaction understanding, temporal reasoning, frame and clip selection, multimodal representation learning, prompting, semantic alignment, and model adaptation. Particular attention is given to graph-based and object-centric reasoning as mechanisms for modeling relations among hands, objects, actions, and scene context over time. We further examine how first-person perception and multimodal foundation models support wearable assistance, robot skill learning, human-to-robot transfer, and embodied decision making. Across the reviewed literature, a consistent limitation emerges: current models recognize visible objects more reliably than evolving interactions, actions, and user intent, especially over long activities. We therefore identify temporally grounded reasoning, interaction-aware supervision, efficient long-video processing, multimodal fusion, graph-enhanced representations, cross-domain generalization, privacy, and trustworthy evaluation as key priorities for deployable embodied intelligence.