中文生成式搜索引擎引用和呈现了什么?一项大规模实证研究
What Do Chinese-Language Generative Search Engines Cite and Surface? A Large-Scale Empirical Study
浏览论文内容
中文总结 AI 辅助
该研究对四个主流平台的中文生成式搜索进行大规模实证,构建数据集并分析引用行为等。发现品牌在答案中选择性呈现,内容契合度等在预测模型中重要,不同时效性查询拟合半衰期不同及界面间来源集有差异,揭示了中文生成式搜索系统特点及界面类型的重要性。
中文摘要 AI 辅助
生成式人工智能问答系统越来越多地在信息获取中起中介作用,将内容可见性从排名搜索结果转移到生成答案中的检索、引用和呈现。我们对四个主流平台的网络和应用程序界面上的中文生成式搜索进行了大规模实证研究。受控设计涵盖八个平台界面、614个查询,每个查询-平台-界面组合进行三次重复。从214,119条原始记录中,我们构建了一个包含160,860条记录的清理后的引用级数据集,并分析了引用行为、来源归因、实体曝光和跨界面一致性。有五个发现。首先,引用库中的品牌在答案中被选择性呈现:总体品牌选择率为8.3%,12.4%包含联系信息的检索来源为答案提供了联系信息。其次,内容契合度、跨来源出现次数和语义角色在预测模型中相对重要,而5118-百度综合质量得分不是任何考察结果的主要预测因素。第三,在有发布日期的被引用页面中,高时效性查询的拟合半衰期约为39天,低时效性查询的拟合半衰期约为68天。第四,约13%的品牌曝光无法与同期引用库匹配,约71%的联系信息曝光无法与爬取的正文匹配。第五,同一平台的应用程序和网络界面之间的来源集存在系统性差异。这些结果描述了中文生成式搜索系统如何选择、归因和呈现信息,并表明界面类型是一个重要的分析维度。
英文摘要
Generative AI question-answering systems increasingly mediate information access, shifting content visibility from ranked search results to retrieval, citation, and presentation in generated answers. We conduct a large-scale empirical study of Chinese-language generative search across the Web and App interfaces of four mainstream platforms. The controlled design covers eight platform interfaces, 614 queries, and three replications per query-platform-interface combination. From 214,119 raw records, we construct a cleaned citation-level dataset of 160,860 records and analyze citation behavior, source attribution, entity exposure, and cross-interface consistency. Five findings emerge. First, brands in the citation pool were selectively surfaced in answers: the overall brand-selection rate was 8.3%, and 12.4% of retrieved sources containing contact information contributed contact information to answers. Second, content fit, cross-source occurrence count, and semantic role were relatively important in predictive models, whereas the 5118-Baidu Composite Quality Score was not the leading predictor for any examined outcome. Third, among cited pages with publication dates, fitted half-lives were approximately 39 days for high-timeliness queries and 68 days for low-timeliness queries. Fourth, approximately 13% of brand exposures could not be matched to the contemporaneous citation pool, and approximately 71% of contact-information exposures could not be matched to the crawled body text. Fifth, source sets differed systematically between the App and Web interfaces of the same platform. These results characterize how Chinese-language generative search systems select, attribute, and surface information and show that interface type is an important dimension of analysis.