面向任务的对话系统中通过无监督微调学习推理与使用工具
Learning to Reason and Use Tools through Unsupervised Fine-Tuning in Task-Oriented Dialog Systems
- HiTZ Center - Ixa, University of the Basque Country UPV/EHU(巴斯克大学UPV/EHU HiTZ中心 - Ixa)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对任务导向对话系统的动态信息检索缺陷,提出适配ReAct框架的无监督微调方法,使LLMs可访问外部知识,其微调后的8B模型性能超越70B上下文系统,在SIMMC数据集上表现优异。
AI中文摘要:
当前对话系统在动态信息检索方面存在不足,常导致幻觉现象并降低响应准确性。我们通过将ReAct框架适配到任务导向对话场景来解决该问题,使大语言模型(LLMs)能够访问外部知识并生成符合事实的响应。主要而言,我们提出一种无监督微调流水线,通过上下文学习推理轨迹来收集数据;使用基于LLM的评判器过滤高质量样本,以构建可靠的训练集。该流水线还通过无监督自我改进循环得到增强,其中改进后的检查点会生成越来越好的轨迹,用于后续微调迭代。在SIMMC数据集上的实验表明,基于ReAct的系统因具备更出色的推理与工具使用能力,性能优于基线系统。值得注意的是,我们微调后的8B模型超越了70B上下文系统。最后,我们还开展了错误分析、场景复杂度的影响分析以及跨领域泛化研究。
英文摘要:
Current dialogue systems struggle with dynamic information retrieval, often leading to hallucinations and lower response accuracy. We address this by adapting the ReAct framework for Task-Oriented Dialogue, enabling Large Language Models (LLMs) to access external knowledge and produce factual responses. Mainly, we propose an unsupervised fine-tuning pipeline that harvests reasoning trajectories via in-context learning inference. High-quality samples are filtered using an LLM-based judge to construct a robust training set. This is enhanced by a unsupervised self-improvement loop, where improved checkpoints generate increasingly better trajectories for subsequent fine-tuning iterations. Experiments on the SIMMC dataset demonstrate that ReAct-based systems outperform baselines due to superior reasoning and tool use. Notably, our fine-tuned 8B model surpasses a 70B in-context system. Finally, we present an error analysis, impact of scene complexity, and cross-domain generalization.