arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁会被提及:引用类型可预测基础语言模型的个体提及情况,而一份名册工具仅捕捉到其中的0.5%

Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It

Dmitrij Żatuchin

arXiv 2607.23893首次发表:更新:

发表机构

Estonian Entrepreneurship University of Applied Sciences (EUAS)(爱沙尼亚应用科学创业大学(EUAS))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在买家选人的类别中,通过API调用研究语言模型提及个体情况,发现类别主导、模型有差异,引用类型可预测提及,引用量不能,还通过名册测量个体AI可见性,发现其只覆盖模型行为的一小部分且不具代表性。

AI 中文摘要

先前关于人工智能品牌知名度的研究衡量的是公司层面:模型是否推荐一家公司,以及这是否与该公司声誉相符。本研究将问题细化到买家选择个人的类别。在2026年7月24日的两个小时内进行了2400次基于实际情况的API调用:120个买家意向提示,四种模型(GPT - 5.6 Sol、Gemini 3.6 Flash、Perplexity Sonar Pro、Grok 4.5),每种模型进行五次迭代,涉及四个欧洲市场和五种查询语言。对每个回复是否提及个体专业人员进行编码,采用一种从不参考名册且排除解析为同名美国城市的检测的规则级联(精确率96.9%,召回率61.7%,因此以下所有比率均为下限)。所有推理都对提示内的聚类进行了校正:组内相关系数0.258,有效样本量407(名义样本量为2400)。模型在25.8%的回复中提及了个体。类别起主导作用:房地产领域为35.4%,汽车经销商领域为32.9%,而保险领域为9.1%(卡方检验值159.3,校正后p = 5.8e - 8)。模型之间差异达四倍,从Grok的38.0%到Gemini的9.3%。引用类型可预测提及情况,而引用量则不能:提及回复中引用个体自身网站的频率高2.6个百分点(95%置信区间为+1.4至+3.9),引用类别门户的频率高4.3个百分点,引用公司自有页面的频率相同(分别为44.1%和45.5%)。在九对匹配的翻译对中,英语提示在36.7%的回复中提及了个体,而当地语言相同问题的提及率为15.6%(优势比3.14,聚类p = 0.074,所以方向明确但设计无法确定)。通过公开的领英搜索构建的939人名单与27293个名字形状的提及中的128个匹配(0.47%),939人中的26人曾被提及,名单得出的0.0%至25.4%的比率衡量了这种重叠情况。基于名单的个体人工智能可见性测量只看到了模型行为中一小部分且不具代表性的情况。

英文摘要

Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the question one level down, in categories where the buyer picks a person. It issued 2,400 grounded API calls in one two-hour window on 24 July 2026: 120 buyer-intent prompts, four models (GPT-5.6 Sol, Gemini 3.6 Flash, Perplexity Sonar Pro, Grok 4.5), five iterations each, four European markets and five query languages. Every response was coded for whether it named an individual professional, by a rule cascade that never consults a roster and that drops detections resolving to a same-named American city (precision 96.9%, recall 61.7%, so every rate below is a lower bound). All inference corrects for clustering within prompt: intraclass correlation 0.258, effective n 407 against a nominal 2,400. Models named an individual in 25.8% of responses. Category dominates: real estate 35.4% and car dealerships 32.9% against insurance 9.1% (chi-square 159.3, p = 5.8e-8 after correction). Models differ four-fold, from Grok 38.0% to Gemini 9.3%. Citation type predicts naming and citation volume does not: naming responses cite the individual's own site 2.6 points more often (95% CI +1.4 to +3.9) and category portals 4.3 points more often, and cite firm-owned pages at the same rate (44.1% against 45.5%). On nine matched translation pairs, English prompts named an individual in 36.7% of responses against 15.6% for the same question in the local language (OR 3.14, clustered p = 0.074, so the direction is clear and the design cannot close it). A 939-person roster built from public LinkedIn search matched 128 of 27,293 name-shaped mentions (0.47%), 26 of the 939 people were ever named, and the roster-derived rates of 0.0% to 25.4% measure that overlap. Roster-based measurement of individual AI visibility sees a small and unrepresentative slice of what models do.

Comments28 pages, 5 figures. Data: https://doi.org/10.5281/zenodo.21612690

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑