arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过结构化专家判断对多语言模型系统进行不确定性感知信任估计

Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement

Jiawei Zheng, Jiazhen Zhang

arXiv 2607.20529首次发表:更新:

发表机构

DigitLab, University of Exeter(数字实验室,埃克塞特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多语言模型聚合问题,核心方法是采用结构化专家判断,通过上下文感知校准问题及库克式对数加权估计专家可靠性,主要贡献是在异构和受污染情况下实现更好的准确性-可靠性平衡,强调聚合时需校准信任。

AI 中文摘要

大语言模型集成越来越多地用于通过组合多个大语言模型的预测来提高可靠性。然而,现有的聚合方法通常假设所有模型同样值得信赖,忽略了不确定性质量的差异。这种假设不适用于异构大语言模型,其可靠性和能力差异很大,使得简单聚合容易受到不可靠或对抗性专家的影响。在这项工作中,我们将多语言模型聚合表述为一个不确定性感知信任估计问题。我们采用决策理论中的结构化专家判断,使用上下文感知校准问题根据概率预测的质量来估计专家的可靠性。具体来说,我们采用库克式对数加权,惩罚过度自信的错误预测并青睐校准良好的专家。我们在MMLU和MMLU-Pro上对均匀、异构和受污染的专家小组评估了我们的方法。结果表明,虽然聚合方法在均匀设置下表现相似,但库克加权在异构和受污染情况下变得至关重要。它实现了卓越的准确性-可靠性平衡,并且在引入不可靠专家时仍然稳健。这些发现表明,多语言模型聚合不仅需要组合预测,还需要在不确定性下校准信任。

英文摘要

Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, existing aggregation methods typically assume that all models are equally trustworthy, overlooking differences in uncertainty quality. This assumption is poorly suited to heterogeneous LLMs, whose reliability and capability vary significantly, making naive aggregation vulnerable to unreliable or adversarial experts. In this work, we formulate multi-LLM aggregation as a problem of uncertainty-aware trust estimation. We adapt structured expert judgment from decision theory, using context-aware calibration questions to estimate expert reliability based on the quality of its probabilistic predictions. Specifically, we employ Cooke-style log weighting, which penalises overconfident incorrect predictions and favours well-calibrated experts. We evaluate our approach on MMLU and MMLU-Pro across homogeneous, heterogeneous, and contaminated expert panels. Results show that while aggregation methods perform similarly in homogeneous settings, Cooke weighting becomes critical under heterogeneity and contamination. It achieves a superior accuracy-reliability balance and remains robust when unreliable experts are introduced. These findings suggest that Multi-LLM aggregation requires not just combining predictions, but calibrating trust under uncertainty.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑