AI 中文总结
本研究综述多轮对话式AI的发展现状,梳理其在数据、模型等方面的进展,指出多模态交互能力提升但存在记忆等短板,并提出相关研究议程。
AI 中文摘要
对话式AI正从孤立的文本提示转向持续的多模态交互。在真实对话中,用户会明确目标、修改请求、打断回应、切换话题并引入新证据,同时期望系统能在多轮对话中保留上下文,这使得多轮对话成为一项独特挑战,要求系统维护和更新记忆、在跨模态、工具及外部知识间建立回应依据,并适配不同语言和文化。本研究综述了多轮对话式AI,涵盖仅文本对话、AudioLLMs(音频大语言模型)及原生语音系统、多模态与全模态系统、工具增强智能体等类别,围绕数据集与基准、建模范式、训练策略、评估设置及跨领域挑战组织文献。分析显示,多模态支持的进展快于维持会话内连贯交互的能力;尽管系统在跨模态感知、说话与行动能力增强,但当前系统仍在持续记忆、跨轮依据、全双工交互、鲁棒评估及文化适配方面存在困难。最后,本研究提出了针对可在多轮、多模态及不同文化间实现记忆、修改、依据、说话、倾听、行动与适配的系统的研究议程。
英文摘要
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
CommentsMulti-turn Conversational AI; Multimodal Dialogue; AudioLLMs; Conversational Memory; Tool-Augmented Agents; Dialogue Evaluation