arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

2026年GENEA挑战赛:基于Seamless Interaction数据集的语音驱动手势生成的大规模解耦评估

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

Rajmund Nagy, Silvia Arellano García, Hendric Voss, Mihail Tsakov, Taras Kucherenko, Youngwoo Yoon, Gustav Eje Henter

arXiv 2608.10839首次发表:更新:

发表机构

KTH Royal Institute of Technology; Bielefeld University; National Library of Sweden; Electronics and Telecommunications Research Institute (ETRI); Motorica AB(瑞典皇家理工学院; 比勒费尔德大学; 瑞典国家图书馆; 电子通信研究院; Motorica AB公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究公布第四届GENEA挑战赛结果,对5个语音驱动手势生成系统开展4项大规模用户研究,发现现有系统在运动真实感、语音对齐等方面仍有不足,相关数据将公开供后续研究。

AI 中文摘要

本预印本展示了第四届GENEA挑战赛的结果,该挑战赛是对参赛团队在Seamless Interaction双人对话数据集上训练的五个语音驱动手势生成系统进行的大规模人类评估。与2023年GENEA挑战赛类似,我们采用了解耦评估方法,在不混淆运动质量和语音对齐两者的情况下对其进行评估,并开展了双人不匹配研究以分离倾听和回应对话者的效果。我们还引入了一项新的语义手势生成任务,以及使用数据的Grounded Gestures子集的文本不匹配评估方法。总共开展了四项大规模用户研究,从869名测试参与者处收集了超过23000张投票。在运动真实感研究中,数据集的过滤片段的运动质量显著高于所有挑战赛提交结果(成对胜率为68%-95%)。在语音对齐研究中,动作捕捉片段提供了62%对齐分数的概念上限,而排名最高的提交结果显著落后,仅为32%,其余提交结果仅略高于与输入无关系统预期的0%。在双人研究中,动作捕捉再次设定了65%适宜性分数的上限,但没有任何提交结果的得分显著高于随机水平,表明这些系统目前还无法对对话者做出回应。最后,语义不匹配评估发现数据集中的手势具有高度表现力(测试参与者在79%的情况下识别出匹配的文本转录),但几乎所有提交结果都未能生成具有语义表现力的动作,最佳结果仅达到8%的适宜性分数。收集到的投票和输出将在该httpsURL公开提供,以促进可复现性和进一步研究。

英文摘要

This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the interlocutor. We additionally introduce a new semantic gesture-generation task and a text-mismatching evaluation methodology using the Grounded Gestures subset of the data. In total, we ran four large-scale user studies, collecting over 23,000 votes from 869 test-takers. In the motion-realism study, the dataset's filtered segments had substantially higher motion quality than all challenge submissions (68-95% pairwise winrate). In the speech-alignment study, the motion-capture segments provided a conceptual ceiling at 62% alignment score, with the top submission significantly behind at 32% and the rest only slightly above the 0% expected of an input-independent system. In the dyadic study, motion capture again set the ceiling at 65% appropriateness score, but no submission scored substantially above chance, indicating that the systems could not yet respond to the interlocutor. Finally, the semantic mismatching evaluation found highly expressive gestures in the dataset (test-takers identified the matching transcript 79% of the time), yet almost all submissions failed to generate semantically expressive motion, with the best achieving only an 8% appropriateness score. The collected votes and outputs will be made publicly available at https://genea-workshop.github.io/2026/challenge/ to facilitate reproducibility and further research.

Comments15 pages, 14 figures. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑