发表机构
Estonian Entrepreneurship University of Applied Sciences(爱沙尼亚创业应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究测试LLM品牌推荐审计中的系统归因,发现单条响应可高准确率识别系统(97.84%),但聚合品牌画像跨领域不可迁移,且未交叉设计无法分离系统与领域或工具。
AI 中文摘要
对AI可见性的审计通常将已部署语言模型的品牌推荐汇总为每个系统的画像。我们测试了这样的画像是否能在包含6,475条存储响应(其中6,324条可分析)的语料库上描述系统,该语料库收集自2025年12月至2026年2月,来自五个已部署端点,涉及礼品推荐、企业声誉和类别归属查询。收集工具截断了许多答案:在1,024个token的输出上限下,Gemini 3 Flash在类别归属查询中有83.1%的答案在句子中途结束。将所有答案截断至前800个字符后,一个字符n-gram分类器(通过提示属性进行交叉验证)将单条响应归因于GPT-5.2、Gemini 3 Flash、带搜索的Gemini 3 Flash、Grok或Perplexity sonar-pro,准确率达97.84%(5,028条响应,383个提示,多数类占比31.5%,30个分割种子)。仅使用长度特征时准确率降至多数类比率,24个格式统计特征达到95.79%,而屏蔽品牌名称和大写标记后准确率仍为97.72%。在保留的查询条件下,按规模加权准确率保持97.43%,未加权为88.0%;在改变收集工具的检索增强臂中,没有Grok答案被归因于Grok(0/120)。将数据聚合为50个模型-领域-条件单元后,十二个行为特征在分组交叉验证下以66.53%的准确率区分四个系统,而标签置换零分布的平均值为33.71%,第95百分位为46.0%。跨领域时,聚合画像失效:在类别归属单元上训练的森林将所有22个礼品单元归因于错误的系统,这与品牌数量的反转一致(礼品中每条响应8.41个品牌对比0.94个,类别归属中3.01对比3.91),而单条响应以89.92%的平衡准确率迁移。答案的表面形式在测试的查询领域中携带了系统信息;聚合的品牌行为则不能,且未交叉的设计无法将系统与领域或收集工具分离。
英文摘要
Audits of AI visibility summarise the brand recommendations of deployed language models into per-system profiles. We test whether such a profile describes the system on one corpus of 6,475 stored responses (6,324 analysable) collected between December 2025 and February 2026 from five deployed endpoints across gift-recommendation, corporate-reputation and category-ownership queries. The collection harness cut many answers short: 83.1% of Gemini 3 Flash answers in category ownership end mid-sentence under a 1,024-token output cap. With every answer cut to its first 800 characters, a character n-gram classifier cross-validated by prompt attributes one response to GPT-5.2, Gemini 3 Flash, Gemini 3 Flash with search, Grok or Perplexity sonar-pro with 97.84% accuracy (5,028 responses, 383 prompts, majority class 31.5%, 30 split seeds). Length alone falls to the majority rate, 24 formatting statistics reach 95.79%, and masking brand names and capitalised tokens leaves 97.72%. Held-out query conditions keep 97.43% weighted by size and 88.0% unweighted; in a retrieval-grounded arm that changes the harness, no Grok answer is attributed to Grok (0/120). Aggregated into 50 model-by-domain-by-condition units, twelve behavioural features separate four systems at 66.53% under grouped cross-validation, against a label-permutation null with mean 33.71% and 95th percentile 46.0%. Across domains the aggregate profile fails: a forest trained on category-ownership units assigns all 22 gift units to the wrong system, consistent with a reversal in brand volume (8.41 against 0.94 brands per response in gifts, 3.01 against 3.91 in category ownership), while single responses transfer at 89.92% balanced accuracy. The surface form of an answer carries the system across the query domains tested; aggregated brand behaviour does not, and the uncrossed design cannot separate the system from the domain or the harness.
Comments30 pages, 5 figures, 9 tables. Appendix D documents corrections to an earlier manuscript