发表机构
Kyoto University; Sony Computer Science Laboratories(京都大学; 索尼计算机科学实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将口头反馈式回应与头部点头整合到配备语音克隆及LLM响应的AI克隆中,经35人被试内研究证实,添加互动倾听行为可提升AI克隆的感知专注度、真实感与共同在场感,其保真度需涵盖倾听行为。
AI 中文摘要
模仿特定人物的AI克隆通常会复制该人物的说话内容和声音,但不会复制其倾听方式。本研究调查添加多模态倾听行为是否能提升此类克隆的存在感与真实性。我们将由实时预测模型驱动的口头反馈式回应(backchannels)和头部点头动作,整合到配备语音克隆与基于大语言模型(LLM)响应的AI克隆中。在一项被试内研究(N=35)中,添加这些行为显著提升了虚拟形象的感知专注度、与真实人物交谈的感觉以及共同在场感。这些结果表明,AI克隆的保真度应超越语音和响应内容,延伸至互动式倾听行为。
英文摘要
AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen. We investigate whether adding multimodal listening behaviors gives such a clone more presence and authenticity. We integrated verbal backchannels and head nodding, driven by real-time prediction models, into an AI clone equipped with voice cloning and LLM-based responses. In a within-subjects study (N=35), adding these behaviors significantly improved the perceived attentiveness of the avatar, the sense of talking with the real person, and the feeling of co-presence. These results indicate that AI clone fidelity should extend beyond voice and response content to include interactive listening behavior.
CommentsThis paper has been accepted to the Late-Breaking Results (LBR) track of the 28th International Conference on Multimodal Interaction (ICMI 2026)