arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05708cs.AIcs.LGcs.MA

CUSP:多智能体多模态推理的可分解集体不确定性

CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

Chung-En Johnny Yu, David Garcia, Brian Jalaian, Nathaniel D. Bastian

首次发表
浏览论文内容

中文总结 AI 辅助

提出CUSP框架,通过语义意见池化分解集体不确定性,无需训练即可在多智能体多模态推理中有效检测错误和指导弃权(不执行),提升系统可靠性。

中文摘要 AI 辅助

聚合异构视觉语言模型(VLM)可以提升多模态推理能力,但单个模型的置信度或聚合答案的置信度都无法在系统层面衡量可靠性。我们提出CUSP(通过语义意见池化的集体不确定性),一个无需训练的量化不确定性框架,将多个VLM的响应映射到共享的语义响应空间,将其池化为池化语义意见,并报告两个互补的系统级信号:集体不确定性(池化意见的离散度)和Jensen-Shannon散度(JSD,模型级意见之间的冲突)。在这个池化语义意见中,未归一化的集体熵精确分解为各模型个体语义熵的均值与JSD之和,将总离散度与模型冲突分离开来。CUSP既不需要token对数概率,也不需要校准标签,适用于开放权重和商业VLM。在静态多VLM集成中,集体不确定性在小模型场景下是最强信号(预测错误检测AUROC为0.764,弃权(不执行)AUARC为0.889),比不确定性基线多数投票和朴素选择高出4.7到15.8个百分点,并且随着集成规模增大优势扩大;JSD在评估的商业场景中最强(AUROC为0.819,AUARC为0.910),并以最高0.982的AUROC对困难答案的模型冲突进行排序。池化预测相比平均单个模型准确率提升5.6到13.0个百分点。在多步骤多智能体系统的完整轨迹中,子智能体集体不确定性以0.619的AUROC将系统故障排序高于随机水平,并在评估的信号中给出最佳弃权(不执行)排序(AUARC为0.699)。

英文摘要

Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an individual model's confidence nor that of the aggregated answer measures reliability at the system level. We present CUSP (Collective Uncertainty through Semantic Opinion Pooling), a training-free uncertainty quantification framework that maps multiple VLM responses to a shared semantic response space, pools them into a pooled semantic opinion, and reports two complementary system-level signals: collective uncertainty, the dispersion of the pooled opinion, and Jensen-Shannon divergence (JSD), the conflict among the model-level opinions. Within this pooled semantic opinion, the unnormalized collective entropy decomposes exactly into the mean of the models' individual semantic entropies and the JSD, separating total dispersion from model conflict. Requiring neither token logits nor calibration labels, CUSP applies to open-weight and commercial VLMs alike. In static multi-VLM ensembles, collective uncertainty is the strongest signal in the small-model regime (0.764 AUROC for prediction-error detection, 0.889 AUARC for abstention), outperforming uncertainty baselines majority voting and naive selection by 4.7 to 15.8 points and widening its margin as the ensemble grows; JSD is strongest in the evaluated commercial regime (0.819 AUROC, 0.910 AUARC) and ranks hard-answer model conflict with AUROC up to 0.982. The pooled prediction also improves accuracy over the average single model by 5.6 to 13.0 points. Over the full trajectory of a multi-step, multi-agent system, subagent collective uncertainty ranks system failures above chance (0.619 AUROC) and gives the best abstention ordering among the evaluated signals (0.699 AUARC).

发表机构

  • University of West Florida(西佛罗里达大学)
  • United States Military Academy(美国军事学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑