发表机构
Pine AI; University of Washington(松树人工智能公司; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究用NameRank衡量大语言模型对实体的识别度,通过对多实体多模型探测及独立评判获取分数,发现识别关注可索引工件,奥运资质与知名奖项情况不同,独立创作者中工具与创造者排名有别等,并指出文献计量法难测识别度等结论。
AI 中文摘要
前沿模型在任何检索步骤之前,从自身权重中回忆起的关于个人或工具的信息,往往会塑造人类看到的首个描述,这使得参数化语料库的存在成为一个测量问题。引用能解释约三分之一模型是否识别研究人员的情况;我们针对剩余部分构建了NameRank,这是一个[0,1]的识别分数。对54个群组中的4685个实体,用一个开放式问题在36个模型上进行探测,由独立评判员根据精心策划的黄金标准给出二元判定。合成空实体得分接近零,判定追踪实体而非模型。研究发现:识别关注的是有名称、可索引的工件,而非资质或头衔。每个奥运式资质都低于在职研究人员基线,在知名奖项层面排名反转。对于独立创作者,工具的排名高于其创造者,传播的资质是命名方法或获奖论文。作为著名工件的众多署名贡献者之一,所得认可几乎为零。没有文献计量法能很好地预测识别度;高引用密度机构在相同引用量下比同行识别度更高;在258个新闻事件中,识别取决于峰值显著性而非持续性。自我报告探测显示自我反省读取的是语料库而非自身知识。
英文摘要
What a frontier model recalls about a person or tool from its own weights -- before any retrieval step -- often shapes the first description a human sees, making that parametric corpus presence a measurement problem. Citations explain about a third of whether a model recognizes a researcher; we target the residual and build NameRank, a [0,1] recognition score: each of 4,685 entities in 54 cohorts is probed with one open-ended question across 36 models, and an independent judge returns a binary verdict against a curated gold -- did the model state a specific, non-guessable fact about this exact entity? -- so hallucination, context echo, and guesses earn nothing. Synthetic-null entities hold the floor near zero, and verdicts track the entity, not the model. One thesis organizes the findings: recognition is paid to named, indexable artifacts, not to credentials or titles. Every Olympic-style credential sits below a working-researcher baseline, because no named artifact ships with the medal, yet the ranking inverts at the marquee tier, where Nobel, Turing, and Fields laureates saturate the panel. For independent creators the tool out-ranks its maker, and the credential that does propagate is a named method or awarded paper. Being one of many named contributors to a celebrated artifact, by contrast, earns almost nothing -- the authors listed on a flagship model report or system card sit near the recognition floor -- because recognition attaches to the artifact's own distinctive name, not to the roster behind it. No bibliometric predicts recognition well; top-density institutions out-recognize peers at matched citations; and on 258 news events recognition loads on peak salience, not persistence. A self-report probe shows introspection reads a corpus prior, not its own knowledge.