arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

鲁棒一致性共识:基于共形预测的多智能体LLM作为评审的区间评估

Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction

Lihui Liu

arXiv 2609.06367首次发表:更新:

发表机构

Wayne State University(韦恩州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种基于共形预测的多智能体LLM作为评审评估的鲁棒不确定性估计框架,通过构建多个LLM评分的预测区间,实现更稳定可靠的评估结果。

AI 中文摘要

LLM作为评审(LLM-as-a-Judge)已成为评估自然语言生成的一种有前景的范式。然而,此类评估相关的不确定性在很大程度上仍未得到探索,这限制了其在实际应用中的可靠性。尽管共形预测(conformal prediction)为不确定性量化提供了原则性框架,但现有方法通常将其应用于单个LLM评审,忽视了使用不同LLM评估器所带来的变异性。在本工作中,我们提出了一种针对多智能体LLM作为评审评估的鲁棒不确定性估计框架。我们的方法为来自多个LLM的基于LLM的评分构建共形预测区间。通过考虑来自不同LLM评审的区间,我们获得了更稳定和可靠的不确定性估计。大量实验表明,我们的方法产生了具有覆盖保证的有效预测区间,并且跨多个评审的基于区间的聚合导致了更稳定的评估结果。

英文摘要

LLM-as-a-Judge has emerged as a promising paradigm for evaluating natural language generation. However, the uncertainty associated with such evaluations remains largely unexplored, which limits their reliability in real-world applications. Although conformal prediction offers a principled framework for uncertainty quantification, existing approaches typically apply it to a single LLM judge, overlooking the variability introduced by using different LLM evaluators. In this work, we propose a robust uncertainty estimation framework for multi-agent LLM-as-a-Judge evaluation. Our approach constructs conformal prediction intervals for LLM-based scores from multiple LLMs. By considering intervals from different LLM judges, we obtain more stable and reliable uncertainty estimates. Extensive experiments demonstrate that our method produces valid prediction intervals with coverage guarantees, and that interval-based aggregation across multiple judges leads to more stable evaluation outcomes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑