arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

盲人摸象:探究长尾分歧知识下大语言模型(LLM)的认知近视问题

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

Zhuoshi Pan, Junru Lu, Yan Qian, H. Vicky Zhao, Di Yin, Xing Sun

arXiv 2608.28478首次发表:更新:

发表机构

Tsinghua University; Tencent; University of Warwick(清华大学; 腾讯; 华威大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出ElephantBench探针,发现LLM在长尾分歧知识问答中仅52.4%的问题能同时回忆两种解释,模型规模扩大等方法无法消除认知不完整性,为评估LLM认知严谨性提供了工具。

AI 中文摘要

事实型问答(QA)通常假设存在唯一的标准答案,却忽略了大语言模型(LLM)是否保留了长尾事实的不同解释。为解决这一缺口,我们推出ElephantBench,这是一个闭卷知识探针,包含1094个通过可审计的基于图的流程生成的问题。该流程从低曝光网络语料库中检索相关文档,识别自然产生的分歧,并将其转换为多解释QA记录。每个答案都经过原始文档和权威公共网络来源的验证,随后由人工标注者审核。在32个模型中,即使是最强的模型也仅在52.4%的问题上同时恢复两种解释,而在几乎所有剩余问题上,它会回忆起一种解释但遗漏另一种。扩大模型规模和推理时推理能提高召回率,但无法消除这种不完整性。语料库分析进一步显示,曝光不平衡有利于主导解释,而少数侧曝光更多则与更完整的召回率相关。这些发现确立了ElephantBench作为可复现的知识探针,用于诊断参数记忆中的认知近视。更广泛地说,我们基于图的基准构建流程提供了一种高效且可扩展的方式,将长尾语料库转换为可溯源的知识探针,支持评估和推进下一代LLM的认知严谨性。代码可在该https URL获取。

英文摘要

Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long-tail facts. To address this gap, we introduce ElephantBench, a closed-book knowledge probe comprising 1,094 questions generated through an auditable graph-based pipeline. The pipeline retrieves related documents from a low-exposure web corpus, identifies naturally occurring disagreements, and converts them into multi-account QA records. Each answer is verified against the originating documents and authoritative public web sources and is then reviewed by human annotators. Across 32 models, even the strongest model recovers both accounts on only 52.4% of questions, while on nearly all remaining questions it recalls one account but omits the other. Scaling model size and inference-time reasoning improve recall but do not eliminate this incompleteness. Corpus analysis further shows that exposure imbalance favors the dominant account, whereas greater minority-side exposure is associated with more complete recall. These findings establish ElephantBench as a reproducible knowledge probe for diagnosing epistemic myopia in parametric memory. More broadly, our graph-based benchmark construction pipeline provides an efficient and scalable way to turn long-tail corpora into source-traceable knowledge probes, supporting efforts to evaluate and advance the epistemic rigour of next-generation LLMs. Code is available at https://github.com/Tencent/ElephantBench.

Comments10 pages, 10 figurs, 1 table, under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑