arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23086cs.AI

POOL:基于相似样本传播的不确定性

POOL: Propagated Uncertainty Over Lookalikes

Rounak Sharma, Ananya B. Sai, Soumyabrata Pal

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出高性价比置信度估计框架POOL,结合混合估计器Hy@5,在多个数据集和黑盒大语言模型上,以更少采样数实现更高AUROC,同时大幅节省生成成本。

中文摘要 AI 辅助

黑盒大语言模型需要置信度得分以区分可能正确与可能错误的输出,这能让系统优先安排人工审核、将不确定案例路由至更强模型,或在开发数据上选择弃权(不执行)阈值。然而现有置信度估计器面临成本-质量权衡:口头置信度成本低但常过度自信,而基于采样的不确定性信息更丰富,但随每个查询的采样数线性扩展。我们提出POOL(全称Propagated Uncertainty Over Lookalikes),这是一种借鉴分组测试思想的高性价比框架,可解决上述权衡问题。POOL将具有重叠的查询词干聚类,在代表性 medoids(质心)上评估基础估计器,将置信度得分软传播至邻近查询,并选择性评估高分歧案例。我们用Hy@$p$实例化该框架,Hy@$p$是结合口头置信度与采样答案的负冯·诺依曼熵计算的光谱答案多样性的混合估计器。在三个数据集的六个领域和五个黑盒大语言模型上,Hy@5的平均AUROC高于口头置信度和Vn@10采样,且使用的采样数仅为Vn@10的一半。POOL-Hy@5保留了其93.5%至97.9%的AUROC,同时节省了19.3%至39.3%的生成成本;在释义密集的工作负载上,生成节省率升至73%至76%,表明可利用语义冗余降低置信度估计成本。

英文摘要

Black-box large language models need confidence scores that can separate likely-correct from likely-incorrect outputs, enabling systems to prioritize human review, route uncertain cases to stronger models, or choose abstention thresholds on development data. Yet existing confidence estimators face a cost-quality trade-off: verbal confidence is cheap but is often overconfident, while sampling-based uncertainty is more informative but scales linearly with the number of samples per query. We propose \textsc{POOL} (\emph{Propagated Uncertainty Over Lookalikes}),a cost-efficient framework that addresses this trade-off taking inspiration from group-testing.\textsc{POOL} clusters query stems with overlaps, evaluates a base estimator on representative medoids, softly propagates confidence scores to nearby queries, and selectively evaluates high-disagreement cases. We instantiate this framework with \textsc{Hy@}$p$, a hybrid estimator that combines verbal confidence with spectral answer diversity computed from the negative von Neumann entropy of sampled answer embeddings.Across six domains from three datasets and five black-box LLMs, \textsc{Hy@}5 achieves higher average AUROC than verbal confidence and \textsc{Vn@}10 sampling while using half as many samples as \textsc{Vn@}10. \textsc{POOL}-\textsc{Hy@}5 retains 93.5--97.9\% of its AUROC while saving 19.3--39.3\% of generations. On paraphrase-dense workloads, generation savings rise to 73-76\%, showing that semantic redundancy can be leveraged to lower confidence-estimation costs.

发表机构

  • Adobe Research, India(印度奥多比研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑