发表机构
Zhejiang University; Beijing Jiaotong University(浙江大学; 北京交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DIGHT框架,耦合数字人交互生成器与人形跟踪策略,通过物理解耦扩散DPO和力反馈偏好,提升交互生成的物理合理性与保真度。
AI 中文摘要
近期方法在生成两个数字人之间的交互方面取得了有前景的进展,这些方法主要依赖基于物理的跟踪策略将数字参考动作转换为可执行的轨迹。然而,有限的跟踪能力限制了可成功执行的参考动作范围,降低了数据利用率。此外,即使跟踪成功,也不能保证物理上合理的响应或忠实实现预期交互。在本文中,我们介绍了DIGHT,一个协同自适应框架,将数字人交互生成器与人形跟踪策略耦合。我们的DIGHT首先使用固定跟踪器在仿真中执行多个文本条件交互候选。然后,它从产生的rollout中构建物理接地偏好,涵盖一般可执行性和交互保真度。我们没有将这些信号压缩为用于候选排序的单一标量奖励,而是使用物理解耦扩散直接偏好优化(DPO)对齐预训练生成器,保留标准特定监督而无需通过模拟器进行微分。为了提高可执行性,偏好对从跟踪误差、摩擦和漂浮中导出。此外,为了提高交互保真度,我们提出将模拟器的力反馈作为接触保真度的度量,并构建关于接触发生、位置、持续时间和力大小的偏好。对齐后的生成器为微调跟踪器提供参考动作,提高生成与物理执行之间的兼容性。大量实验表明,我们的方法不仅提高了生成动作的物理合理性,还使仿真中的人形交互更加可靠和忠实。
英文摘要
Recent methods have made promising progress in generating interactions between two humanoids, largely relying on physics-based tracking policies to convert digital reference motions into executable trajectories. However, limited tracking capabilities restrict the range of reference motions that can be successfully executed, reducing data utilization. Moreover, even successful tracking does not guarantee physically plausible responses or faithful realization of the intended interactions. In this paper, we introduce DIGHT, a co-adaptive framework that couples a Digital human Interaction Generator with a Humanoid Tracking policy. Our DIGHT first executes multiple text-conditioned interaction candidates in simulation using a fixed tracker. It then constructs physics-grounded preferences from the resulting rollouts, covering both general executability and interaction fidelity. Rather than collapsing these signals into a single scalar reward for candidate ranking, we align the pretrained generator using physics-decoupled diffusion direct preference optimization (DPO), preserving criterion-specific supervision without differentiating through the simulator. To improve executability, preference pairs are derived from tracking error, friction, and floating. Additionally, to improve interaction fidelity, we propose to incorporate force feedback from simulator as a measure of contact fidelity and construct preferences over contact occurrence, location, duration, and force magnitude. The aligned generator then supplies reference motions for fine-tuning the tracker, improving compatibility between generation and physical execution. Extensive experiments demonstrate that our approach not only improves the physical plausibility of generated motions but also enables more reliable and faithful humanoid interactions in simulation.