LLMPEDIA:浏览、验证与对比大语言模型的参数化百科知识
LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs
浏览论文内容
中文总结 AI 辅助
该研究开发了LLMPEDIA工具,可测量并浏览大语言模型参数化百科知识的偏差,通过审计断言事实性,发现其真实率低于基准测试,提供多种视图供用户查看对比
中文摘要 AI 辅助
旗舰语言模型在MMLU等基准测试中看似已趋于饱和,得分超过90%——但基准测试仅测试实验者想到要问的内容,存在固定问题集的可得性偏差。LLMPEDIA可使这种偏差可测量且可浏览。我们从三个模型家族(GPT-5-mini、DeepSeek-V3.2、Llama-3.3-70B)的参数化记忆中递归生成了约130万篇文章,未使用检索技术,随后对分层抽样的原子断言对照维基百科和精选网络栈进行审计,将每个断言标记为已支持、已反驳或证据不足(Saeed和Razniewski,2026)。在均匀随机抽样中,真实率为68.4%——比MMLU低21个百分点以上,其中30.5%的断言证据不足:这些断言是任何基准测试都未探测到的,且世界上最大的百科全书也无法裁决的长尾知识或看似合理的幻觉,证据无法区分——将GPTKB为三元组建立的覆盖缺口扩展到了自由文本领域(Hu等,2025)。由此产生的实时开放百科全书让访问者可通过五种一键视图逐一检查这一前沿内容:链接遍历探索、断言层面的事实性、跨模型与政治角色对比,以及引导式主题下钻——每个页面、断言和裁决都有稳定的URL。LLMPEDIA已上线,访问地址为该httpsURL
英文摘要
Flagship language models appear saturated on benchmarks like MMLU (Hendrycks et al., 2021), scoring above 90% - yet benchmarks test only what the experimenter thought to ask, the availability bias of fixed question sets. LLMPEDIA makes this bias measurable and browsable. We recursively materialized ~1.3M articles from three model families' parametric memory (GPT-5-mini, DeepSeek-V3.2, Llama-3.3-70B) without retrieval, then audited a stratified sample of atomic claims against Wikipedia and a curated web stack, coloring every claim supported, refuted, or insufficient (Saeed and Razniewski, 2026). On a uniform random sample the true rate is 68.4% - more than 21 pp below MMLU - with 30.5% of claims insufficient: assertions no benchmark probes and the world's largest encyclopedia cannot adjudicate - long-tail knowledge or plausible hallucination, the evidence cannot tell - extending to free text the coverage gap GPTKB established for triples (Hu et al., 2025). The resulting live, open encyclopedia lets visitors inspect this frontier one claim at a time through five one-click views - link-traversal exploration, claim-level factuality, cross-model and political-persona comparison, and a guided topic drill-down - each page, claim, and verdict at a stable URL. LLMPEDIA is live at https://llmpedia.net
发表机构
- ScaDS.AI Dresden/Leipzig(ScaDS.AI德累斯顿/莱比锡)
- TU Dresden(德累斯顿工业大学)
机构由 AI 辅助整理,请以论文原文为准。