发表机构
Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出三智能体框架,用于评估大语言模型的问题澄清能力,通过三个不同LLM智能体及多维度指标开展评估,为提升对话式LLM应用的澄清能力提供了结构化方法。
AI 中文摘要
大语言模型(LLMs)正越来越多地部署在交互式系统中,精准理解用户意图至关重要。这类系统的一项关键能力是有效的问题澄清,尤其是在用户查询模糊或信息不足时。本文提出一种新颖的三智能体框架,用于稳健评估LLM开展澄清对话的能力。该框架包含三个不同的基于LLM的智能体:(1)问题澄清智能体(QCA,待评估系统),负责识别歧义并提出澄清问题;(2)应答智能体(RA),旨在模拟人类用户的回复,可能包含不相关或具有挑战性的回答;(3)评估智能体(EA,一种LLM作为评判者),基于一套综合指标评估对话质量。本文以供应链领域为例,详细介绍了合成数据生成的方法,提出了评估歧义处理、问题质量、对话效率、语言恰当性及最终意图对齐的指标,还简要讨论了EA针对人类判断的验证。本工作为基准测试、验证和提升对话式LLM应用的澄清能力提供了结构化方法。
英文摘要
Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarification, especially when user queries are ambiguous or underspecified. This paper introduces a novel tri-agent framework for the robust evaluation of an LLM's ability to engage in clarifying dialogue. Our framework comprises three distinct LLM-based agents: (1) a Question Clarifying Agent (QCA), the system under evaluation, tasked with identifying ambiguities and posing clarifying questions; (2) a Respondent Agent (RA), designed to simulate human user responses, potentially including irrelevant or challenging replies; and (3) an Evaluator Agent (EA), an LLM-as-a-judge, which assesses the quality of the dialogue based on a comprehensive set of metrics. We detail a methodology for synthetic data generation in the supply chain domain as an example. We propose metrics evaluating ambiguity handling, question quality, dialogue efficiency, language appropriateness, and final intent alignment. We also briefly discuss the validation of the EA against human judgments. This work provides a structured approach to benchmark, validate, and improve the clarification capabilities of conversational LLM applications.