arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.09798cs.CV

测试时自体-外体适应用于动作预见的多标签原型生长与双线索一致性

Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency

  • University of Electronic Science and Technology of China(电子科学与技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu, Fanman Meng, Lili Pan, Hongliang Li

更新

AI总结:

本文提出TE²A任务,通过多标签原型生长模块和双线索一致性模块实现测试时自体-外体适应,提升动作预见性能,实验表明优于现有方法。

AI中文摘要:

高效的自体(Ego)与外体(Exo)视图适应对于人机协作等应用至关重要。然而,大多数现有自体-外体适应方法依赖目标视图数据进行训练,从而增加计算和数据收集成本。本文首次探索了测试时自体-外体适应用于动作预见(TE²A)任务,旨在在测试时调整已训练的源视图模型以预测目标视图动作。现有测试时适应(TTA)方法难以解决此任务,因为存在多动作候选和显著的时空视图间隙。为此,我们提出了一种增强的双线索原型生长网络(DCPGN),通过积累多标签知识并整合跨模态线索实现有效的测试时自体-外体适应和动作预见。具体而言,我们提出多标签原型生长模块(ML-PGM)通过多标签分配和基于置信度的再加权平衡多个正类,通过熵优先队列策略更新。然后,双线索一致性模块(DCCM)引入轻量级叙述者生成文本线索指示动作进程,补充包含各种物体的视觉线索。此外,我们约束推断的文本和视觉logits以构建双线索一致性,以在时间和空间上连接自体和外体视图。在新提出的EgoMe-anti和现有EgoExoLearn基准上进行了大量实验,结果表明我们的方法有效,显著优于相关最先进方法。代码可在https://github.com/ZhaofengSHI/DCPGN获得。

英文摘要:

Efficient adaptation between Egocentric (Ego) and Exocentric (Exo) views is crucial for applications such as human-robot cooperation. However, the success of most existing Ego-Exo adaptation methods relies heavily on target-view data for training, thereby increasing computational and data collection costs. In this paper, we make the first exploration of a Test-time Ego-Exo Adaptation for Action Anticipation (TE$^{2}$A$^{3}$) task, which aims to adjust the source-view-trained model online during test time to anticipate target-view actions. It is challenging for existing Test-Time Adaptation (TTA) methods to address this task due to the multi-action candidates and significant temporal-spatial inter-view gap. Hence, we propose a novel Dual-Clue enhanced Prototype Growing Network (DCPGN), which accumulates multi-label knowledge and integrates cross-modality clues for effective test-time Ego-Exo adaptation and action anticipation. Specifically, we propose a Multi-Label Prototype Growing Module (ML-PGM) to balance multiple positive classes via multi-label assignment and confidence-based reweighting for class-wise memory banks, which are updated by an entropy priority queue strategy. Then, the Dual-Clue Consistency Module (DCCM) introduces a lightweight narrator to generate textual clues indicating action progressions, which complement the visual clues containing various objects. Moreover, we constrain the inferred textual and visual logits to construct dual-clue consistency for temporally and spatially bridging Ego and Exo views. Extensive experiments on the newly proposed EgoMe-anti and the existing EgoExoLearn benchmarks show the effectiveness of our method, which outperforms related state-of-the-art methods by a large margin. Code is available at \href{https://github.com/ZhaofengSHI/DCPGN}{https://github.com/ZhaofengSHI/DCPGN}.

补充信息

↑