arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20157cs.CV

G3Ego:用于自我中心动作理解的注视引导图

G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding

Marko Haralović, Akash Ramakrishnan, Estefania Talavera Martinez

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出G3Ego框架,将注视直接融入图构建以识别动作相关实体,在EGTEA Gaze+和MECCANO数据集上实现有竞争力的动作理解性能,且在类别不平衡评估下提升Macro-F1,无需昂贵视频预训练。

中文摘要 AI 辅助

自我中心动作理解通常使用在大量异中心数据集上预训练的大型视频模型来解决,但许多第一人称动作依赖于仅涉及少数相关实体的少量手-物交互。我们提出G3Ego,一种用于自我中心动作理解的基于图的框架,它使用注视作为结构线索来识别场景中与动作相关的实体。从稀疏采样的帧中,G3Ego根据视觉-语言描述、已定位的物体和手部线索构建动作场景图,然后利用佩戴相机者的注视修剪不相关的实体。所得的图嵌入被时间聚合以用于动作识别和预测。与之前主要将注视用作辅助模态或注意力信号的工作不同,G3Ego将注视直接融入图的构建中,产生专注于动作相关交互的高效且可解释的表示。在EGTEA Gaze+和MECCANO上的实验表明,G3Ego与基于视频的方法相比达到了有竞争力的性能,并且在类别不平衡评估下持续提高了Macro-F1,同时避免了对计算成本高昂的视频预训练的依赖。这些结果证明了注视引导图表示在自我中心动作理解中的有效性。

英文摘要

Egocentric action understanding is often addressed using large video models pretrained on extensive exocentric datasets. However, many first-person actions depend on a small number of hand-object interactions involving only a few relevant entities. We propose G3Ego, a graph-based framework for egocentric action understanding that uses gaze as a structural cue to identify action-relevant entities in the scene. From sparsely sampled frames, G3Ego constructs action scene graphs from vision-language descriptions, grounded objects, and hand cues, and then prunes irrelevant entities using the camera wearer's gaze. The resulting graph embeddings are temporally aggregated for action recognition and anticipation. Unlike prior work that uses gaze primarily as an auxiliary modality or attention signal, G3Ego incorporates gaze directly into graph construction, producing efficient and interpretable representations focused on action-relevant interactions. Experiments on EGTEA Gaze+ and MECCANO show that G3Ego achieves competitive performance compared with video-based approaches and consistently improves Macro-F1 under class-imbalanced evaluation, while avoiding reliance on computationally expensive video pretraining. These results demonstrate the effectiveness of gaze-guided graph representations for egocentric action understanding.

发表机构

  • University of Zagreb(萨格勒布大学)
  • University of Twente(特文特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑