利用耦合流补全进行人-物交互建模
Harnessing Coupled Stream Completion For Human-Object Interaction Modeling
AI总结:
提出TRACE框架,通过分离并耦合身体、物体和手部运动流的连续潜在表示,实现文本驱动的HOI生成与补全,并在InterAct等基准上取得最优接触指标。
AI中文摘要:
文本条件的人-物交互(HOI)生成要求身体运动、物体轨迹与旋转以及手部关节运动保持协调。这些组成部分在尺度和动态特性上各不相同,但必须在接触、相对姿态和时序上达成一致。共享表示可能会限制每个流各自独特的结构,而独立生成则使每个流无法响应其他流的变化。仅靠潜在空间监督也无法在解码后直接约束接触。我们提出TRACE,一个连续潜在框架,它保持流状态分离并耦合其更新。TRACE将身体、物体和手部运动编码为独立的潜在表示,并根据完整的当前交互状态预测每个流的速度。对解码后运动的几何损失进一步约束了接触和随时间变化的物体相对运动。同一模型支持从其他两个流补全任意缺失的单个流。冻结的流特征也可作为语言模型的输入,用于HOI理解。在InterAct、OMOMO和BEHAVE上的实验表明,联合补全训练改善了生成效果,且冻结的流特征比原始运动编码更能提升理解性能。在InterAct上,TRACE在比较的方法中实现了最高的接触精确率、召回率和F1分数。
英文摘要:
Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories & rotations, and hand articulation to remain coordinated. These components differ in scale and dynamics, but must agree on contact, relative pose, and timing. A shared representation may limit the distinct structure of each stream, while independent generation prevents each stream from responding to changes in the others. Latent supervision alone also does not directly constrain contact after decoding. We propose TRACE, a continuous latent framework that keeps stream states separate and couples their updates. TRACE encodes body, object, and hand motion into separate latents and predicts each stream velocity from the complete current interaction state. Geometric losses on decoded motion further constrain contact and object-relative motion over time. The same model supports completion of any single absent stream from the other two. Frozen flow features also serve as input to a language model for HOI understanding. Experiments on InterAct, OMOMO, and BEHAVE show that joint completion training improves generation and that frozen flow features improve understanding over raw-motion encoding. On InterAct, TRACE achieves the highest contact precision, recall, and F1 among the compared methods.