arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12062cs.CLcs.AI

偏好树优化:通过前瞻模拟增强面向目标的对话

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations

Lior Baruch, Moshe Butman, Kfir Bar, Doron Friedman

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出偏好树优化(PTO)框架,结合带前瞻的偏好树与直接偏好优化(DPO),在动机访谈领域通过虚拟患者等模拟对话生成偏好数据,训练的对话智能体在关键指标上优于基线。

中文摘要 AI 辅助

开发能够进行多轮面向目标对话的系统仍是一项重大挑战,尤其是在数据有限的专业领域。本研究提出一种名为偏好树优化(Preference Tree Optimization, PTO)的新型框架,旨在通过一种名为带前瞻的偏好树的方法生成偏好数据,从而迭代改进此类对话系统中的智能体模型。本研究聚焦于动机访谈(Motivational Interviewing, MI)——一种旨在促进行为改变的咨询技术,我们利用虚拟患者和神谕评估器模拟对话并生成丰富的偏好数据集。通过将该方法与直接偏好优化(Direct Preference Optimization, DPO)相结合,我们旨在在迭代训练周期中增强智能体的决策能力。所提出的框架解决了数据稀缺问题,推动了面向目标领域中更精细、更有效的对话系统的发展。实验评估表明,PTO框架可提升动机访谈(MI)领域中对话智能体在面向目标对话中的性能。经PTO训练的模型在会话满意度、工作联盟等关键指标上始终优于基线模型。此外,引入前瞻模拟可改善长期规划并提升对话策略的有效性,更深的前瞻配置可产生最稳定、得分最高的结果。

英文摘要

Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree Optimization (PTO), designed to iteratively improve agent models in such dialogue systems, by generating preference data using a method called Preference Tree with Look-Ahead. Focusing on Motivational Interviewing (MI) -- a counseling technique aimed at facilitating behavioral change -- we leverage virtual patients and an oracle evaluator to simulate conversations and generate rich preference datasets. By combining this method with Direct Preference Optimization (DPO), we aim to enhance the agent's decision-making capabilities over iterative training cycles. The proposed framework addresses data scarcity and advances the development of more nuanced and effective dialogue systems in goal-oriented domains. Experimental evaluations demonstrate that the PTO framework enhances dialogue agents' performance in goal-oriented conversations within the domain of Motivational Interviewing (MI). Models trained with PTO consistently outperformed the baseline in key metrics such as session satisfaction and working alliance. Additionally, incorporating look-ahead simulations led to improved long-term planning and more effective conversational strategies, with deeper look-ahead configurations yielding the most stable and high-scoring results.

发表机构

  • Reichman University(赖希曼大学)
  • School of Computer Science(计算机科学学院)
  • School of Communications(传播学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑