审计百度与谷歌AI搜索中的来源曝光
Auditing Source Exposure in Baidu and Google AI Search
查看机构详情
- Politecnico di Milano(米兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究跨语言审计百度与谷歌AI搜索概览,发现平台-语言设置在来源曝光上差异显著且重叠低,强调评估AI搜索需兼顾答案内容与来源可见性分布。
中文摘要 AI 辅助
AI生成的概览正日益成为搜索界面中一个突出的层,但其在中文搜索中的行为仍未得到充分探索。我们使用从MS MARCO采样的英文查询及其翻译后的中文对应查询,对百度和谷歌上的AI概览行为进行了跨语言审计。我们的分析考察了概览在跨平台-语言设置中何时被触发,哪些主域名在中文概览中获得可见曝光,这种曝光的集中程度如何,以及来源重叠在不同设置间的变化。我们还比较了匹配查询意图下生成答案的基于嵌入的语义相似度。结果显示,在概览可用性和可见来源曝光方面,不同平台-语言设置之间存在显著差异。在聚合层面,各设置在可见主域名清单上表现出低重叠度,而匹配查询答案的中位余弦相似度范围从0.701到0.813。这些发现表明,答案层面的语义相似度和聚合来源曝光捕捉了AI中介搜索的不同维度。因此,对AI搜索的评估不仅应考虑生成答案的内容,还应考虑来源可见性如何在平台、语言和信息环境中分布。
英文摘要
AI-generated overviews are becoming an increasingly prominent layer of search interfaces, yet their behavior in Chinese-language search remains underexplored. We conduct a cross-lingual audit of AI overview behavior on Baidu and Google using English queries sampled from MS MARCO and their translated Chinese counterparts. Our analysis examines when overviews are triggered across platform-language settings, which host domains receive visible exposure in Chinese-language overviews, how concentrated that exposure is, and how source overlap varies across settings. We also compare the embedding-based semantic similarity of generated answers for matched query intents. The results reveal substantial differences across platform-language settings in overview availability and visible source exposure. At the aggregate level, the settings exhibit low overlap in visible host-domain inventories, while matched-query answers yield median cosine similarities ranging from 0.701 to 0.813. These findings indicate that answer-level semantic similarity and aggregate source exposure capture distinct dimensions of AI-mediated search. Evaluations of AI search should therefore consider not only the content of generated answers but also how source visibility is distributed across platforms, languages, and information environments.