发表机构
Xidian University; National University of Defense Technology; Human Institute of Advanced Technologyy; Shanghai Road Transport Development Center(西安电子科技大学; 国防科技大学; 人类高级技术研究院; 上海市道路运输发展中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出KDA方法,用LLMs分解动作标签为原子动作,经KIM和KDM解缠,在多标签动作识别基准上达SOTA性能,且模块可通用集成。
AI 中文摘要
复杂场景下的动作识别常涉及多个并发的细粒度动作,对内部动作结构的建模构成挑战。现有多数方法依赖整体表征,不足以捕捉细微交互和细粒度语义。近期基于提示的方法引入了解缠,但缺乏显式语义指导,仅基于视觉或结构化线索的方法仍较粗粒度。本文提出基于原子动作的知识引导解缠(KDA),利用细粒度语义知识增强动作表征,实现更精准的解缠。具体而言,使用大语言模型(LLMs)将动作标签分解为原子动作,提供显式时空语义。知识注入模块(KIM)首先将原子动作知识整合到视频特征中。基于该增强表征,知识解缠模块(KDM)进一步解缠原子动作知识,为动作解缠生成更精准的语义指导。引入知识解缠损失(KD Loss)以促进KDM内知识成分的更清晰解缠。大量实验表明,KDA提升了特征判别性,在多标签动作识别基准上达到了最先进性能。此外,KIM和KDM可便捷集成到其他方法中,展现出强通用性。
英文摘要
Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures. Most existing methods rely on holistic representations, which are insufficient for capturing subtle interactions and fine-grained semantics. While recent prompt-based approaches introduce disentanglement, they lack explicit semantic guidance, and methods based solely on visual or structured cues remain coarse-grained. In this paper, we propose Knowledge-guided Disentanglement with Atomic Actions (KDA), which leverages fine-grained semantic knowledge to enhance action representations and enable more precise disentanglement. Specifically, we use Large Language Models (LLMs) to decompose action labels into atomic actions, providing explicit spatial-temporal semantics. A Knowledge Injection Module (KIM) first integrates atomic action knowledge into video features. Based on this enhanced representation, a Knowledge Disentanglement Module (KDM) further disentangles atomic action knowledge to produce more precise semantic guidance for action disentanglement. A Knowledge Disentanglement Loss (KD Loss) is introduced to encourage clearer disentanglement of knowledge components within KDM. Extensive experiments demonstrate that KDA improves feature discriminability and achieves state-of-the-art performance on multi-label action recognition benchmarks. Moreover, KIM and KDM can be readily integrated into other methods, demonstrating strong generality.
CommentsACMMM 26