校准用于人类与AI对话的LLM评判器
Calibrating LLM Judges for Human and AI Conversations
浏览论文内容
中文总结 AI 辅助
本研究提出锚定集与校准函数,将LLM评判器得分统一到共享尺度,并发布VA数据集,发现校准可迁移至人机对话,缩小与人类判别差距。
中文摘要 AI 辅助
衡量一段对话的成功程度仍然困难,即使对于评判口语对话的人类来说也是如此。我们在CANDOR数据集上评估了最先进的LLM作为逐点和成对对话成功评判器的表现,发现逐点评分与人类评分的相关性中等,而成对比较则受到长转录文本和位置偏差的影响。由于这导致评判器得分在不同模型间不可比较,我们提出了一个小型锚定集和一个校准函数,该函数可将任何评判器校准到一个共享的、可解释的尺度上。我们进一步发布了Voice Arena目标数据集(VA),包含200个面向任务的人机对话和人类与智能体对话,并带有成对标注,揭示了当前评判器与人类水平判别能力之间的显著差距。利用VA,我们测试了CANDOR拟合的校准是否可迁移到人类与AI对话中,发现尽管在拟合过程中从未见过VA,该校准仍能将评判器置于共享尺度上。
英文摘要
Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of conversational success on CANDOR, finding pointwise scoring correlates moderately with human ratings, while pairwise comparison suffers from long transcripts and positional bias. Since this leaves judge scores incomparable across models, we propose a small anchor set and a calibration function that calibrates any judge onto a shared, interpretable scale. We further release the Voice Arena Goal Dataset (VA), 200 task-oriented human-AI and human-agent conversations with pairwise annotations, revealing a substantial gap between current judges and human-level discrimination. Using VA, we test whether CANDOR-fitted calibration transfers to human-AI conversations, finding it brings judges onto a shared scale despite never observing VA during fitting.
发表机构
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
- Charles University(查理大学)
- ETH Zurich(苏黎世联邦理工学院)
- KTH Royal Institute of Technology(瑞典皇家理工学院)
- NICT(日本信息通信研究机构)
- Voice Arena(语音竞技场)
- University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。