AI 中文总结
该研究针对现有对话评估框架的不足,推出UPHELD基准,开发混合评判者框架,提升了对话评估与人类判断的相关性,为LLM对话评估提供了稳健基础。
AI 中文摘要
随着大语言模型(LLMs)越来越多地被用于开放式多轮交互服务,在人类规模上评估对话质量已成为核心挑战。现有针对摘要、翻译或短格式问答任务构建的评估框架,无法充分衡量人类规模对话的一致性,尤其是这些指标的推导和验证往往依赖合成数据而非人类来源。我们通过引入UPHELD(UPwork人类规模评估长对话)填补这一空白,这是一个用于评估人类规模对话能力的大型、参考完备的基准,超越了事实正确性。UPHELD包含数百个由专业编剧创作的完整人机对话,具有真实的轮次密度,在30000多个专家生成的对话轮次中,每轮有36000多个人类标注。我们使用UPHELD系统评估经典自动指标和无参考的LLM作为评判者方法,发现它们与专家人类判断的相关性不可靠。基于此分析,我们使用UPHELD开发了一种混合评判者(Mixture-of-Judges)框架,该框架结合了多种评估信号,使与人类评估的相关性提高了约30%。总体而言,UPHELD为评估人类规模的对话智能提供了一个稳健、以人类为基础的基础,填补了现有LLM数据集领域的关键空白。
英文摘要
As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for summarization, translation, or short-form QA tasks fall short of adequately measuring the consistency of human-scale dialogue, especially when derivation and validation of these metrics themselves often rely on synthetic rather than human sources. We fill the gap by introducing UPHELD (UPwork Human-Scale Evaluated Long Dialogues), a large, reference-full benchmark for evaluating human-scale conversational ability beyond factual correctness. UPHELD consists of hundreds of complete human-to-human dialogues authored by professional script writers, with realistic turn densities and 36,000+ per-turn human annotations across 30,000+ expert-generated dialogue turns. Using UPHELD, we systematically evaluate classical automatic metrics and reference-free LLM-as-a-judge approaches, and find them unreliable when correlated with expert human judgment. Building off this analysis, we use UPHELD to develop a Mixture-of-Judges framework that combines multiple evaluative signals and improves correlation with human assessments by approximately 30%. Overall, UPHELD provides a robust, human-grounded foundation for evaluating human-scale conversational intelligence that fills a crucial gap in the pre-existing LLM dataset landscape.
CommentsInternational Conference on Machine Learning 2026