发表机构
KAIST; University of Illinois Urbana-Champaign; LG AI Research(韩国科学技术院; 伊利诺伊大学厄巴纳-香槟分校; LG人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有跨视角视频表示学习方法中视角不变与可变语义纠缠的问题,提出PRISM模型,通过语义隐分解与语言监督实现解耦,在多个数据集上达到SOTA,零样本性能优于领域内模型。
AI 中文摘要
跨视角视频表示学习旨在捕捉视角不变的动作语义,尽管自我中心视角与外部视角视频之间存在显著外观变化。然而,现有方法将每个视频编码为统一嵌入,视角不变语义与视角可变语义在共存时不可避免地纠缠——我们表明,即使是明确为视角不变性训练的跨视角方法也存在这种失效模式。我们的核心见解是:当视角不变特征能与任意视角可变特征充分重组且保留各自独立语义时,该特征才真正实现解耦。基于此,我们提出PRISM,它将视频分解为视角不变和视角可变隐变量,并在语言监督下重组这些隐变量,以促进两个流的清晰分解。PRISM在EgoExo4D、EgoExoLearn、AE2上取得了最先进的结果,在零样本设置下甚至超越了领域内模型。代码可在该https URL获取。
英文摘要
Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentric and exocentric videos. However, existing methods encode each video as a unified embedding, where view-invariant and view-variant semantics inevitably entangle under co-occurrences - a failure mode we show persists even in cross-view methods explicitly trained for view-invariance. Our key insight is that a view-invariant feature is truly disentangled when it can be sufficiently recomposed with an arbitrary view-variant feature while preserving their independent semantics. Building on this, we propose PRISM, that decomposes video into view-invariant and view-variant latents and recompose them under language supervision encouraging clean decomposition of the two streams. PRISM achieves state-of-the-art results on EgoExo4D, EgoExoLearn, AE2, even surpassing in-domain models under zero-shot setting. Code is available at https://github.com/litcoderr/prism.