arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越更多视角:面向熟练度估计的冗余感知自我-外部视角融合

Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation

Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert

arXiv 2608.25736首次发表:更新:

发表机构

University of Surrey; University of Wollongong(萨里大学; 卧龙岗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对EgoExo熟练度估计中多视角冗余与过拟合问题,提出AdaMVS与VIB-GB模块,在EgoExo-4D等数据集上取得新SOTA结果。

AI 中文摘要

自我-外部视角(EgoExo)熟练度估计旨在通过整合来自第一人称(自我)视角的细粒度运动线索与多个第三人称(外部)视角的空间上下文,来评估动作质量。简单增加外部视角会降低EgoExo性能,因为冗余或有噪声的视角会稀释有用的运动线索。我们的分析确定了两个关键原因:(1)多视角冗余——从数据角度看,某些视角提供有限或有噪声的信息,稀释了判别性线索;(2)过拟合——从特征角度看,传统融合增加了表示复杂度,导致模型记忆特定视角模式而非学习可泛化的表示。为解决这些问题,我们提出两个互补模块:AdaMVS,从数据角度在弱监督下自适应识别并融合最具信息性的视角标记;VIB-GB,从特征角度结合梯度混合(Gradient Blending)与变分信息瓶颈(Variational Information Bottleneck)正则化,以压缩冗余信号并抑制训练期间的过拟合。在EgoExo-4D和EgoExo-Fitness数据集上的实验表明,我们的方法既学会关注哪些视角,又学会如何融合它们,取得了新的最优结果。我们的源代码可在该https URL获取。

英文摘要

EgoExo proficiency estimation aims to assess action quality by integrating fine-grained motion cues from egocentric (1st-person) views with spatial context from multiple exocentric (3rd-person) views. Simply adding more exocentric views degrades EgoExo performance, as redundant or noisy perspectives dilute useful motion cues. Our analysis identifies two key causes: (1) Multiview redundancy - From the data perspective, certain views provide limited or noisy information, diluting discriminative cues; (2) Overfitting - From the feature perspective, conventional fusion increases representational complexity, causing the model to memorise view-specific patterns rather than learn generalisable representations. To address these issues, we propose two complementary modules: AdaMVS, which adaptively identifies and fuses the most informative view tokens under weak supervision from the data perspective, and VIB-GB, which combines Gradient Blending and Variational Information Bottleneck regularisation from the feature perspective to compress redundant signals and suppress overfitting during training. Experiments on EgoExo-4D and EgoExo-Fitness demonstrate that our method learns both which view to look at and how to fuse them, achieving new state-of-the-art results. Our source code is available at https://github.com/dx199771/AdaMVS

CommentsAccepted by ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑