发表机构
School of Artificial Intelligence, Nanjing University of Science & Technology; School of Intelligence Science and Technology, Nanjing University(南京理工大学人工智能学院; 南京大学智能科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态咨询师回复生成中推理与回复不一致的问题,提出MOCC语料库和两阶段框架MOCC-R1,通过监督微调与强化学习优化一致性,实验验证其有效性。
AI 中文摘要
多模态咨询师回复生成(MCRG)旨在根据多模态对话历史生成恰当的咨询师回复。该领域的进展受到两个差距的限制:首先,现有数据集很少包含由合格咨询师进行的、持续的人类记录咨询互动;其次,现有方法未显式优化咨询推理与生成回复之间的一致性,这可能削弱MCRG系统的可靠性。为此,我们引入了MOCC,一个多模态咨询对话语料库,包含超过200小时的互动,涉及154名经资质认证的咨询师。基于MOCC,我们提出了MOCC-R1,一个用于优化推理-回复一致性的两阶段框架。冷启动监督微调训练模型生成结构化轨迹,该轨迹包括来访者状态理解、将咨询原则与计划行动相联系的回复意图,以及最终回复。随后,强化学习(RL)奖励基于计划的连贯性和计划执行,鼓励推断的状态和计划得到对话上下文支持,并鼓励回复实现该计划。实验证明了所提出的MOCC-R1的有效性。
英文摘要
Multimodal counselor response generation (MCRG) aims to generate an appropriate counselor response from multimodal dialogue histories. Progress is limited by two gaps: first, existing datasets rarely capture sustained, human-recorded counseling interactions conducted by qualified counselors; Second, existing methods do not explicitly optimize consistency between counseling reasoning and the generated response, potentially undermining the reliability of MCRG systems. Thus, we introduce MOCC, a multimodal counseling conversation corpus containing over 200 hours of interactions involving 154 credential-verified counselors. Based on MOCC, we propose MOCC-R1, a two-stage framework for optimizing reasoning-response consistency. Cold-start supervised fine-tuning trains the model to generate a structured trajectory consisting of client-state understanding, a response intent that links a counseling principle to a planned action, and the final response. Reinforcement learning (RL) then rewards grounded plan coherence and plan execution, encouraging the inferred state and plan to be supported by the dialogue context and the response to realize that plan. Experiments demonstrate the effectiveness of the proposed MOCC-R1.