arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REALM:一种用于具身反应式聆听的从粗到细生成框架

REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

Peizhen Li, Longbing Cao, Yang Zhang

arXiv 2609.33095首次发表:更新:

AI 中文总结

REALM提出一种从粗到细的音频驱动反应式聆听框架,通过门控融合与随机残差细化,在ViCo和L2L上提升运动质量,并成功部署于Ameca机器人。

AI 中文摘要

生成响应式听者面部运动是具身对话式AI的一项重要任务。两个建模挑战是核心:在保持与听者持续运动连续性的同时,考虑说话者提示的时间;以及在整体运动轨迹中捕捉局部可变的面部事件。听者反应可能在前述提示之后出现时间滞后,而短暂的表情和眨眼引入了难以确定性预测的变化。这些挑战促使了一个结合历史感知时间对齐与随机表情细化的框架。我们提出REALM(反应式具身音频驱动聆听模型),一种用于音频驱动反应式聆听的从粗到细框架。一个反应式门控说话者-听者融合模块通过延迟中心注意力先验和自适应门控,将听者运动历史与说话者音频相结合。一个粗解码器预测基础运动轨迹,该轨迹在表情子空间中通过音频条件随机残差进行增强,同时保留粗姿态参数。在ViCo和L2L上的评估显示,在多个运动质量指标上相对于评估基线有所改进。额外的分析检查了延迟敏感性、门控行为和眨眼动态。最后,在Ameca人形机器人上的部署和感知用户研究证明了生成行为对物理具身的适用性。代码:此https URL 演示:此https URL

英文摘要

Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ

Comments22 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑