arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11446cs.AI

校准感知的不确定性级联用于高效异构模型协作

Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration

Yilin Zhang, Han Jiang, Cai Xu, Ying Liu, Wei Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

提出校准感知的不确定性级联(CAUC),通过独立校准模型置信度并统一决策标准,实现异构模型高效协作,在语言和图像基准上分别提升准确率1.9%并减少47%强模型调用、降低GFLOPs达57%。

中文摘要 AI 辅助

异构模型协作旨在利用不同模型的互补优势,以平衡预测性能和推理成本。现有方法通常依赖于训练好的路由器(其将路由决策绑定到固定的任务和模型池)或原始置信度级联(其阈值在异构模型间缺乏一致的可信度语义)。因此,这些方法难以适应变化的模型池和部署预算。我们提出了校准感知的不确定性级联(CAUC),一种简单的后处理框架,它独立校准每个模型的置信度,并使用验证数据选择部署策略。由此产生的校准置信度分数建立了一个共同的可信度尺度,用于接受早期预测、调用更强的模型或选择性地组合模型输出。这一统一的决策标准将部署策略与任何特定的模型池或运营预算解耦。我们进一步从理论上证明,校准赋予置信度阈值一种明确的择风险解释,而未校准的分数则无法提供类似的可信度保证。大量实验表明,在六个语言基准上,CAUC相对于仅使用强模型的推理实现了平均1.9%的相对准确率提升,同时避免了约47%的强模型调用。在图像分类基准上,它在保持或提升预测性能的同时,将测得的GFLOPs最多减少了57%。

英文摘要

Heterogeneous model collaboration seeks to exploit the complementary strengths of different models to balance predictive performance and inference cost. Existing approaches typically rely either on trained routers, which tie routing decisions to a fixed task and model pool, or on raw-confidence cascades, whose thresholds lack consistent reliability semantics across heterogeneous models. Consequently, these approaches adapt poorly to changing model pools and deployment budgets. We propose Calibration-Aware Uncertainty Cascades (CAUC), a simple post-hoc framework that independently calibrates each model's confidence and selects deployment policies using validation data. The resulting calibrated confidence scores establish a common reliability scale for accepting an early prediction, invoking a stronger model, or selectively combining model outputs. This unified decision criterion decouples deployment policies from any particular model pool or operating budget. We further show theoretically that calibration gives confidence thresholds an explicit selective-risk interpretation, whereas uncalibrated scores offer no comparable reliability guarantee. Extensive experiments demonstrate that, across six language benchmarks, CAUC achieves an average relative accuracy improvement of 1.9% over strong-model-only inference while avoiding approximately 47% of strong-model calls. On image classification benchmarks, it maintains or improves predictive performance while reducing measured GFLOPs by up to 57%.

发表机构

  • School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑