Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
语言模型的不确定性量化:一套黑盒、白盒、LLM评判者和集成评分器
机构 * CVS Health(CVS健康公司)
专题命中 评测与基准 :LLM(title,abstract);language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出了一种灵活的框架,利用黑盒、白盒、LLM评判者和集成评分器来检测语言模型的幻觉问题,通过可调节的集成方法提升检测性能。
Comments Accepted by TMLR; UQLM repository: https://github.com/cvs-health/uqlm
Journal ref Transactions on Machine Learning Research, 2025