发表机构
Beth Israel Deaconess Medical Center; Harvard Medical School; Pontificia Universidad Javeriana; Tufts University(贝斯以色列女执事医疗中心; 哈佛医学院; 哈维里亚那天主教大学; 塔夫茨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对ChatGPT、Perplexity、Google AI Overview三款工具的AI心理健康信息查询引用开展多平台多语言审计,发现引用高度集中、平台偏好差异大等问题,发布相关工具供审计生成式健康搜索使用。
AI 中文摘要
在线健康信息查询正从用户考虑排名链接的关键词搜索,转向生成单一答案并整理其引用文献的对话系统。因此,来源评估的责任从用户转移到平台,但这些系统呈现的内容特征却鲜为人知。我们在两种提示条件下,针对20个英文心理健康问题,审计了三款免费消费产品(ChatGPT、Perplexity、Google AI Overview),其中3个问题的子集还被翻译成另外6种不同资源层级的语言。我们记录了1140个响应中的15942条引用文献,涉及1713个唯一域名,随后用一个由确定性分类器应用的九类组织类型对每条引用进行分类,该分类器已通过人工编码验证。引用文献高度集中:被引用最多的10个域名占英文引用的43.6%,政府、商业健康和学术来源的占比相近,均约为22%。各平台在典型引用量上差异不大,但在一致性和偏好的来源类型上差异显著。明确要求提供来源对内容构成的影响仅较为轻微。非英文查询呈现的引用更少,且被导向语言适配资源的比例显著更低。我们将该类型、分类器和带注释的语料库作为可重复使用的工具发布,用于审计生成式健康搜索。
英文摘要
Online health information seeking is shifting from keyword search, where users consider a ranked list of links, to conversational systems that compose a single answer and curate its citations. Source evaluation therefore passes from user to platform, yet what these systems surface is poorly characterized. We audited three free consumer products (ChatGPT, Perplexity, Google AI Overview) on twenty English mental health questions under two prompt conditions, with a subset of three also translated into six further languages of varying resource tiers. We recorded 15,942 citations across 1,140 responses and 1,713 unique domains, then classified every citation with a nine-category organizational typology applied by a deterministic classifier validated against human coding. Citations were heavily concentrated: the ten most-cited domains accounted for 43.6% of English citations, and government, commercial health, and academic sources were closely matched at roughly 22% each. Platforms differed little in typical citation volume but sharply in consistency and in the source types they favored. Explicitly requesting sources shifted composition only modestly. Non-English queries surfaced fewer citations and were routed to language-appropriate resources at significantly lower rates. We release the typology, classifier, and annotated corpus as reusable instruments for auditing generative health search.
Comments28 pages (16-page main text plus supporting information), 5 figures, 5 tables. Under review. Code: https://github.com/mindbench-ai/search-source-audit Data: https://huggingface.co/datasets/MindBench/search-source-audit