arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2305.15699cs.CV

从外中心到自我中心视角的跨视角动作识别理解

Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective

Thanh-Dat Truong, Khoa Luu

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出一种新颖的跨视角动作识别学习方法(CVAR),通过引入基于几何的约束和跨视角自注意力损失,有效将外中心视频知识转移到自我中心视角,并在多个基准上取得最先进性能。

中文摘要 AI 辅助

理解自我中心视频中的动作识别已成为一个具有众多实际应用的重要研究课题。由于自我中心数据收集规模的限制,学习基于深度学习的鲁棒动作识别模型仍然困难。由于不同视角下视频的差异,将从大规模外中心数据中学到的知识转移到自我中心数据具有挑战性。我们的工作引入了一种新颖的跨视角动作识别学习方法(CVAR),能够有效地将知识从外中心视角转移到自我中心视角。首先,基于分析两个视角间的相机位置,我们向Transformer的自注意力机制中引入了一种新颖的基于几何的约束。然后,我们提出了一种在未配对的跨视角数据上学习的新的跨视角自注意力损失,以强制自注意力机制学习跨视角的知识转移。最后,为了进一步提升跨视角学习方法的性能,我们提出了有效衡量视频和注意力图中相关性的指标。在标准自我中心动作识别基准(即Charades-Ego、EPIC-Kitchens-55和EPIC-Kitchens-100)上的实验结果表明了我们方法的有效性和最先进的性能。

英文摘要

Understanding action recognition in egocentric videos has emerged as a vital research topic with numerous practical applications. With the limitation in the scale of egocentric data collection, learning robust deep learning-based action recognition models remains difficult. Transferring knowledge learned from the large-scale exocentric data to the egocentric data is challenging due to the difference in videos across views. Our work introduces a novel cross-view learning approach to action recognition (CVAR) that effectively transfers knowledge from the exocentric to the selfish view. First, we present a novel geometric-based constraint into the self-attention mechanism in Transformer based on analyzing the camera positions between two views. Then, we propose a new cross-view self-attention loss learned on unpaired cross-view data to enforce the self-attention mechanism learning to transfer knowledge across views. Finally, to further improve the performance of our cross-view learning approach, we present the metrics to measure the correlations in videos and attention maps effectively. Experimental results on standard egocentric action recognition benchmarks, i.e., Charades-Ego, EPIC-Kitchens-55, and EPIC-Kitchens-100, have shown our approach's effectiveness and state-of-the-art performance.

发表机构

  • University of Arkansas(阿肯色大学)

机构由 AI 辅助整理,请以论文原文为准。

↑