arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10008cs.IRcs.CLcs.LG

大语言模型推荐系统是否知道自己在产生幻觉?针对目录保真度的置信度校准审计

Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

Srijith Ravikumar

首次发表
浏览论文内容

中文总结 AI 辅助

该研究审计了四个零样本LLM推荐系统的幻觉率与置信度校准,发现其系统性置信不足,提出conformal弃权阈值效果有限,建议审计需同时报告校准与OOD率并采用锚定目录的提示。

中文摘要 AI 辅助

用于Top-K项目推荐的大语言模型(LLM)推荐系统经常输出目标目录之外的项目标题。以往的审计将此衡量为二元的域外(OOD)率,但从未探究模型是否知晓自己在产生幻觉。我们联合审计了来自四个独立厂商的四个零样本LLM推荐系统(Mistral Large、Llama-3.3-70B、GPT-OSS-120B、Claude Sonnet 4.6,均为未接地或微调的系统)的幻觉率(OOD@10)和口头表达的置信度校准(ECE、Brier、可靠性),覆盖三个目录(MovieLens-25M、Amazon Reviews 2023 Toys、Yelp Open Dataset),并按项目流行度分层。幻觉具有目录依赖性:MovieLens上为0-0.2%,Amazon上为4.5-8.3%,Yelp上为2.2-8.4%;但即使幻觉为0时,口头置信度仍存在显著校准误差(MovieLens上OOD为0%时,ECE最高达0.223)。所有四个LLM在全部12个分组中均系统性地表现为置信不足,对推荐项目的口头评分均值为67-86,而推荐准确率达92-100%,这与LLM幻觉研究中通常强调的过度自信相反。这种置信不足最好被解读为“提示不匹配”:“仅询问”提示会引出通用的推荐质量评分,而非目录成员概率。在口头置信度上应用 conformal 弃权(不执行)阈值,在α∈{0.05,0.10,0.15,0.20}下最多可将幻觉降低0.7个百分点,但会带来4-21个百分点的覆盖率成本:置信不足的通道无法区分正确项目与幻觉,因此该阈值主要移除了正确项目。我们建议,对LLM推荐系统的审计应同时报告校准情况与OOD率,并使用锚定目录的提示而非通用置信度提示。

英文摘要

LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits report a binary out-of-domain rate; none ask whether the model knew. We jointly audit hallucination rate (OOD@10) and verbalized-confidence calibration (ECE, Brier, reliability) for four zero-shot LLM recommenders from four independent vendors (Mistral Large, Llama-3.3-70B, GPT-OSS-120B, Claude Sonnet 4.6), not grounded or fine-tuned systems, across three catalogs (MovieLens-25M, Amazon Reviews 2023 Toys, Yelp Open Dataset), stratified by item popularity. Measuring catalog membership is itself the hard part: on identical outputs the reported rate moves by an order of magnitude with the string matcher used, and F1 cannot separate the candidates. We validate the instrument against 201 human judgments and select on net bias, where the adopted one is off by -0.040 against +0.144 for the common fuzzy rule. Hallucination is then strongly catalog-dependent (0.6-2.7% on MovieLens, 11.6-38.7% on Yelp, 49.3-61.0% on Amazon Toys). Each model holds a near-constant confidence level barely responsive to the catalog, while the catalog-hit rate swings 60 points, so the sign of the error is set by where a model's constant lands against a catalog's accuracy: 7 of the twelve cells are under-confident and 5 over-confident, all four under-confident on MovieLens, all four over-confident on Amazon Toys. We read this as an elicitation mismatch: "Just Ask" elicits a generic quality rating, not a catalog-membership probability. A conformal abstention threshold over verbalized confidence changes hallucination by at most 1.65 pp across four alpha levels, because the channel cannot separate correct items from hallucinations. We recommend that audits report calibration alongside OOD, validate the matcher producing the OOD number, and use catalog-anchored elicitation.

发表机构

  • Amazon.com LLC(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑