HUMAN-TCI:面向文本到运动检索的以躯干为中心的层次化多流运动感知网络
HUMAN-TCI: Hierarchical Multi-Stream Motion-Aware Network with Torso-Centered Interaction for Text-to-Motion Retrieval
浏览论文内容
中文总结 AI 辅助
提出HUMAN-TCI层次化多流运动感知网络,通过三流架构显式建模躯干交互,高效实现文本到运动检索,支持复杂多动作描述。
中文摘要 AI 辅助
准确检索人体运动是文本引导的人体运动建模与合成的关键第一步,因为它能从大型数据集中选择语义相关的序列,并为下游任务提供有依据的参考。从自然语言描述中检索运动仍然具有挑战性,因为句子可以描述多个动作、重叠运动以及身体部位之间复杂的依赖关系。现有方法通常聚焦于简单、单一动作的描述,且通常独立处理身体部位或仅通过拼接特征来处理,而未显式建模躯干运动如何影响其他部位。此外,其处理流程往往依赖计算量大的模型,引入相当大的开销,尤其是在建模较长或更复杂的运动序列时。这限制了判别性运动模式表征的学习,降低了实际应用中的检索准确性、可解释性和效率。为解决这些局限,我们提出了HUMAN-TCI,一种用于文本引导的人体运动检索的层次化多流运动感知网络。HUMAN-TCI采用三流架构,分别建模上半身、下半身和躯干运动,同时显式捕获它们之间的交互,使躯干相关运动能够影响其他身体部位的位置和动态。通过纳入定制的躯干注意力,我们的模型能有效识别复杂的人体运动模式,捕获细粒度的运动关系,并处理复杂的多动作描述。我们的框架支持对简单、单一动作句子以及包含顺序或重叠动作的长篇组合描述的检索,而无需依赖复杂模型。
英文摘要
Accurate retrieval of human motions is a crucial first step in text-guided human motion modeling and synthesis, as it selects semantically relevant sequences from large datasets and provides grounded references for downstream tasks. Retrieving motions from natural language descriptions remains challenging because sentences can describe multiple actions, overlapping movements, and intricate dependencies between body parts. Existing methods often focus on simple, single-action descriptions and typically process body parts independently or by merely concatenating features, without explicitly modeling how torso movements influence other parts. In addition, their processing pipelines often rely on computationally heavy models, introducing considerable overhead, particularly when modeling longer or more complex motion sequences. This limits learning discriminative motion-pattern representations, reducing retrieval accuracy, interpretability, and efficiency in practical applications. To address these limitations, we propose HUMAN-TCI, a Hierarchical Multi-Stream Motion-Aware Network for text-guided human motion retrieval. HUMAN-TCI employs a three-stream architecture that separately models upper-body, lower-body, and torso motions while explicitly capturing their interactions, allowing torso-related movements to influence the positioning and dynamics of other body parts. By incorporating tailored torso attention, our model effectively recognizes complex human motion patterns, captures fine-grained motion relationships and handles complex multi-action descriptions. Our framework supports retrieval for both simple, single-action sentences and long, compositional descriptions containing sequential or overlapping actions without relying on complex models.
发表机构
- James Cook University(詹姆斯·库克大学)
- Macquarie University(麦考瑞大学)
机构由 AI 辅助整理,请以论文原文为准。