基于大语言模型的多轮任务导向对话合成用于真实推理
LLM-Driven Multi-Turn Task-Oriented Dialogue Synthesis for Realistic Reasoning
中文总结 AI 辅助
本文提出基于LLM的多轮任务导向对话合成框架,通过三级优化生成真实推理场景下的对话,提升LLM的逻辑推理能力。
中文摘要 AI 辅助
大型语言模型(LLMs)的推理能力,定义为其根据输入信息分析、推断和做出决策的能力,对于构建智能任务导向对话系统至关重要。然而,现有基准测试并不足以反映现实世界场景的复杂性,这限制了它们在评估和增强LLM推理能力方面的有效性。许多当前的推理数据集过于简单和抽象,往往与现实任务流程、领域约束和操作规则脱节,使难以有效评估LLM的逻辑推理能力。此外,预训练语料库中的数据污染会削弱评估结果的可靠性,而传统的数据集构建众包方法则劳动强度大且难以扩展。为了解决这些挑战,我们提出了一种基于LLM的框架,用于合成基于现实推理场景的多轮任务导向对话,利用三级优化来提升对话质量。我们的方法生成基于真实任务场景的对话,富含现实世界信息,并表现出强上下文连贯性。相应的推理任务围绕这些对话精心设计,并迭代优化以持续提高任务的质量和挑战性。所得到的数据集成为评估和推动LLM真实逻辑推理能力的重要基准。实验结果表明,基于合成数据的推理任务引入了非平凡的推理挑战,并为提高LLM的推理能力提供了有意义的支持。
英文摘要
The reasoning capability of large language models (LLMs), defined as their ability to analyze, infer, and make decisions based on input information, is essential for building intelligent task-oriented dialogue systems. However, existing benchmarks do not sufficiently reflect the complexity of real-world scenarios, which limits their effectiveness in evaluating and enhancing LLM reasoning in practical contexts. Many current reasoning datasets are overly simplistic and abstract, often disconnected from realistic task flows, domain constraints, and operational rules, making it difficult to effectively evaluate LLMs' logical reasoning ability. In addition, data contamination from pretraining corpora undermines the reliability of evaluation results, and traditional crowdsourcing methods for dataset construction are labor-intensive and difficult to scale. To address these challenges, we propose a LLM-driven framework for synthesizing multi-turn, task-oriented dialogues grounded in realistic reasoning scenarios, leveraging trilevel optimization to enhance dialogue quality. Our method generates dialogues grounded in authentic task scenarios, enriched with real-world information, and exhibiting strong contextual coherence. Corresponding reasoning tasks are carefully designed around these dialogues and iteratively refined to continuously improve the tasks' quality and challenge. The resulting dataset serves as a valuable benchmark for assessing and advancing the realistic logical reasoning capabilities of LLMs. Experimental results show that our synthetic data-based reasoning tasks introduce non-trivial reasoning challenges and provide meaningful support for improving the reasoning capabilities of LLMs.