arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26529cs.CL

面向开放式对话中Pairwise LLM评判的多专家共形风险控制

Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue

Ming Cheng, Yusheng Dai, Qiuhong Ke, Zhaolin Chen, Lizhen Qu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出多专家共形风险控制算法,针对开放式对话的Pairwise LLM评判问题,设计Score Averaging、Decision Voting及MC3方法,构建Panel基准数据集,提升了同构与异构专家面板的评估性能。

中文摘要 AI 辅助

本文探索用于开放式对话中Pairwise LLM作为评判者评估的多专家共形风险控制(Conformal Risk Control, CRC)算法。我们的核心见解是,多专家聚合为CRC提供了互补的解决方案:CRC通过弃权(不执行)在决策阈值处控制风险,而聚合则从源头优化评分函数。基于此,我们首先设计两种多专家CRC方法:Score Averaging(分数平均)和Decision Voting(决策投票),分别在分数和决策层面进行聚合。虽然两种策略在同构专家面板上均优于单专家方法,但在异构LLM评判者上,它们仍保持风险有效性,但仅恢复有限的覆盖率,因为统一阈值无法匹配专家不同的评分尺度。为解决此问题,我们进一步提出Marginal-Calibrated Conformal Consensus(MC3,边际校准共形共识):它通过初始阈值比率捕捉每个专家的不同尺度,同时联合调整在校准和测试中统一应用的决策函数$C_t(x)$,从而保持可交换性。为评估我们的框架,我们构建了Panel,这是一个包含1800对的开放式对话人类偏好基准数据集。该数据集基于四个开放权重LLM在三个领域(ESConv、MSC、DREAM)的对话上下文生成的响应构建,且可获取完整logit。实验中,我们发现Score Averaging和Decision Voting在同构面板上显著提升准确率和接受率;值得注意的是,MC3通过适配三个数据集上每个专家的不同评分尺度,将这些提升扩展到异构面板。

英文摘要

In this paper, we explore multi-expert Conformal Risk Control (CRC) algorithms for pairwise LLM-as-a-Judge evaluation in open-ended dialogue. Our core insight is that multi-expert aggregation offers a complementary remedy to CRC: whereas CRC controls risk at the decision threshold through abstention, aggregation sanitizes the scoring function at its source. Guided by this, we first design two multi-expert CRC methods: Score Averaging and Decision Voting, which aggregate at the score and decision levels, respectively. While both strategies outperform single-expert methods on homogeneous expert panels, on heterogeneous LLM judges they remain risk-valid but recover only limited coverage, because a uniform threshold cannot match the experts' distinct scoring scales. To resolve this issue, we further propose Marginal-Calibrated Conformal Consensus (MC3): it captures distinct per-expert scales via initial threshold ratios, while jointly tuning a unified decision function $C_t(x)$ applied identically in both calibration and test, thereby preserving exchangeability. To evaluate our framework, we construct Panel, a 1,800-pair human pairwise-preference benchmark for open-ended dialogue. It is built on responses generated by four open-weight LLMs over dialogue contexts from three domains (ESConv, MSC, DREAM), with full logit access. In experiments, we find that both Score Averaging and Decision Voting substantially improve accuracy and acceptance rate on homogeneous panels. Notably, MC3 extends these gains to heterogeneous panels by accommodating distinct per-expert scoring scales across all three datasets.

发表机构

  • Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑