探索面向任务型对话的 ReAct 提示:洞见与不足
Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings
- University of Lorraine(洛林大学)
- LORIA(洛林计算机科学与应用实验室)
- Charles University(查理大学)
- Orange Innovation(法国电信创新实验室)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文将 ReAct 提示用于任务型对话并在模拟与真实用户中评估,发现其模拟成功率落后于 SOTA,但人类主观满意度更高。
中文摘要 AI 辅助
大型语言模型(LLM)因其在非结构化对话中的出色能力而广受欢迎。利用推理与行动(ReAct)(Yao 等人,2022)等先进提示策略增强 LLM,已在解决传统上需要强化学习的复杂任务方面展现出前景。在本工作中,我们应用 ReAct 策略来引导 LLM 执行任务型对话(TOD)。我们在模拟环境和真实用户两种场景下评估基于 ReAct 的 LLM(ReAct-LLM)。尽管 ReAct-LLM 在模拟中的成功率显著低于最先进方法,但这一差异在人工评估中变得不那么明显。此外,与基线相比,尽管 ReAct-LLM 成功率较低,人类仍报告对其有更高的主观满意度,这很可能得益于其自然且语气自信的回复。
英文摘要
Large language models (LLMs) gained immense popularity due to their impressive capabilities in unstructured conversations. Empowering LLMs with advanced prompting strategies such as reasoning and acting (ReAct) (Yao et al., 2022) has shown promise in solving complex tasks traditionally requiring reinforcement learning. In this work, we apply the ReAct strategy to guide LLMs performing task-oriented dialogue (TOD). We evaluate ReAct-based LLMs (ReAct-LLMs) both in simulation and with real users. While ReAct-LLMs severely underperform state-of-the-art approaches on success rate in simulation, this difference becomes less pronounced in human evaluation. Moreover, compared to the baseline, humans report higher subjective satisfaction with ReAct-LLM despite its lower success rate, most likely thanks to its natural and confidently phrased responses.