arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13876cs.RO

与你共语:机器人共语手势的无训练个性化

Co-Speech with You: Training-Free Personalization of Robot Co-Speech Gestures

  • Tilburg University(蒂尔堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Bosong Ding, Selma Ancel, Giacomo Spigler, Murat Kirtay

AI总结:

提出无训练个性化流程,结合冻结扩散先验与风格编码器,从10秒动作提取风格嵌入,显著提升共语手势个性化(SRA从27.6%升至69.5%),并迁移至实体机器人部署。

AI中文摘要:

个人机器人应能适应新用户的共语手势风格,而无需重新训练模型。我们提出了一种无训练的个性化流程,该流程将冻结的音频条件扩散先验与手势风格编码器及轻量级条件适配器相结合。编码器首先经过训练以区分说话者身份,然后与适配器一起使用扩散目标进行联合优化,从而能够通过单次前向传播从约10秒的注册动作中提取可复用的风格嵌入。为支持此设置,我们还发布了一个Quest 3采集应用程序和一个包含十名参与者自发共语动作的数据集。我们使用风格识别准确率(SRA)和弗雷歇手势距离(FGD)在保留的说话者上评估系统,以衡量个性化和动作质量。我们的方法将SRA从冻结先验的27.6%提高到69.5%,同时保持动作质量(FGD 34.2对比34.8),而用另一个人的注册嵌入替换则使SRA降至11.4%。生成的手势被重定向到实体NAO机器人,并且这种改进也迁移到机器人部署场景中,在该场景中使用五种TTS语音合成语音,保留了67.3%的SRA。

英文摘要:

Personal robots should adapt their co-speech gesture style to a new user without requiring model retraining. We present a training-free personalization pipeline that combines a frozen audio-conditioned diffusion prior with a gesture style encoder and lightweight conditioning adapters. The encoder is first trained to discriminate speaker identities and then jointly refined with the adapters using the diffusion objective, enabling a reusable style embedding to be extracted from approximately 10 seconds of enrollment motion through a single forward pass. To support this setting, we also release a Quest~3 capture application and a dataset of spontaneous co-speech motion from ten participants. We evaluate the system on held-out speakers using Style Recognition Accuracy (SRA) and Fr'echet Gesture Distance (FGD) to measure personalization and motion quality. Our approach improves SRA from 27.6\% for the frozen prior to 69.5\% while preserving motion quality (FGD 34.2 versus 34.8), and replacing the enrollment embedding with another person's reduces SRA to 11.4\%. The generated gestures are retargeted to a physical NAO robot, and this improvement also transfers to the robot deployment setting, where speech is synthesized using five TTS voices, retaining 67.3\% SRA.

↑