arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Inter-X++:用于多模态人与人交互分析的综合基准

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

Liang Xu, Chengqun Yang, Zili Lin, Xintao Lv, Yichao Yan, Xin Jin, Zhibo Chen, Xiaokang Yang, Wenjun Zeng

arXiv 2608.20312首次发表:更新:

发表机构

Shanghai Jiao Tong University; Eastern Institute of Technology; University of Science and Technology of China(上海交通大学; 宁波工程学院; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出多模态人与人交互基准Inter-X++,构建含精细标注的大规模数据集,开发统一HHI框架OpenHHI,其在生成与感知任务上达最优性能,验证了统一表示的有效性。

AI 中文摘要

感知与合成人与人交互(HHI)的能力是开发智能数字人系统的基础。然而,现有数据集和建模方法存在根本局限:运动学保真度低、缺失灵巧手势、多模态标注严重不足;此外,碎片化的交互表示与不一致的评估协议也阻碍了公平且严谨的基准测试。为系统性解决这些瓶颈,我们提出Inter-X++——一个旨在赋能通用HHI分析的综合大规模基准。该基准通过新型混合运动捕捉系统采集,包含11388段高保真交互序列、超810万帧数据,具备精确的全身运动与详细的手指关节运动信息。同时,我们为数据基础补充了多维度标注,包括分层细粒度文本描述、交互类别、因果交互顺序、主体关系与性格,以及顶点级接触图和物理正则化约束。借助这些精细标注,我们构建了涵盖四类下游任务的统一测试平台,该平台对称覆盖生成式与感知式范式。为消除基准歧义,我们系统性标准化了交互表示与评估协议。最后,我们超越数据集构建,提出OpenHHI——一个单一统一的HHI表示与建模框架,可联合优化交互重建与语义理解。大量实验表明,OpenHHI在生成与感知任务上均达到了当前最优性能,这明确证明我们的统一表示可同时衔接交互理解与生成。

英文摘要

The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are fundamentally constrained by low-fidelity kinematics, the omission of dexterous hand gestures and a severe lack of rich multimodal annotations. Furthermore, fragmented interaction representations and inconsistent evaluation protocols also impede fair and rigorous benchmarking. To systematically address these bottlenecks, we present Inter-X++, a comprehensive and large-scale benchmark designed to empower versatile HHI analysis. Captured via a novel hybrid motion capture system, Inter-X++ provides 11,388 high-fidelity interaction sequences and over 8.1M frames, featuring precise whole-body movements and detailed finger articulations. Meanwhile, we enrich the data foundation with multifaceted annotations, including hierarchical fine-grained textual descriptions, interaction categories, causal interaction orders, the relationship and personality of the subjects, as well as vertex-level contact maps and physically regularized constraints. Leveraging these elaborate annotations, we formulate a unified testing ground comprising four categories of downstream tasks that symmetrically span both generative and perceptive paradigms. To eliminate benchmarking ambiguities, we systematically standardize the interaction representations and evaluation protocols. Finally, we go beyond dataset construction to propose OpenHHI, a single and unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments reveal that OpenHHI achieves state-of-the-art performance on both generation and perception tasks. This definitively proves that our unified representation successfully bridges interaction understanding and generation simultaneously.

Comments24 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑