arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22839cs.LGcs.AI

面向黑盒大语言模型分类推理的层级感知监督式不确定性估计

Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning

Shuting Xie, Nathaniel Lesperance, Graham W. Taylor

首次发表
浏览论文内容

中文总结 AI 辅助

针对黑盒LLM在生物多样性监测分类推理中置信度估计难的问题,提出层级感知监督估计器,提升了微AUROC,验证了层级结构对弃权规则的重要性。

中文摘要 AI 辅助

大语言模型(LLMs)越来越多地用于科学决策支持,但在黑盒场景下可靠的置信度估计仍然困难。我们研究长尾生物多样性监测流程中,黑盒LLM生成的层级分类推理的不确定性估计。利用开源工具LLM提取的代理特征,我们训练轻量型监督估计器,采用层级感知监督来预测各分类层级的正确性。在三个工具LLM上,这些监督估计器在微判别和单一全局拒绝阈值下的选择性预测任务中,始终优于基于词元似然的基线,将微AUROC从0.57提升至0.75至0.80。最佳结果由特定层级的多头设计(H3)实现,这表明当需要统一的弃权(不执行)规则时,考虑层级输出结构十分重要。我们的代码可在此URL获取。

英文摘要

Large language models (LLMs) are increasingly used for scientific decision support, yet reliable confidence estimation remains difficult in black-box settings. We study uncertainty estimation for hierarchical taxonomic reasoning generated by a black-box LLM in a long-tailed biodiversity monitoring pipeline. Using proxy features extracted by an open-source tool LLM, we train lightweight supervised estimators with hierarchy-aware supervision to predict rank-wise correctness. Across three tool LLMs, the supervised estimators consistently outperform a token-likelihood baseline for micro discrimination and selective prediction under a single global rejection threshold, improving micro AUROC from 0.57 to 0.75--0.80. The best results are achieved by a rank-specific multi-head design (H3), suggesting that accounting for hierarchical output structure is important when a unified abstention rule is required. Our code is publicly available at https://github.com/uoguelph-mlrg/hierarchy-aware-llm-uq

发表机构

  • Vector Institute for AI(向量人工智能研究所)
  • University of Guelph(圭尔夫大学)

机构由 AI 辅助整理,请以论文原文为准。

↑