发表机构
Heidelberg University(海德堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过随机算法审计发现,LLM辅助医生选择时,声誉信号影响最大,人口统计信号存在倾斜但模型解释未体现,提出重复行为审计是合适的监测技术。
AI 中文摘要
越来越多的患者向大型语言模型(LLM)助手咨询该看哪位医生,这些系统因此成为AI信息中介:算法在人们对他人的选择中起中介作用,从而悄无声息地大规模决定哪些医生能被看到。我们报告一项预先设定的随机算法审计,旨在探究哪些因素会因果性地影响这些推荐。7个模型(6个开源权重模型;gpt-4o-mini)各自在5张合成家庭医生卡片中进行选择,这些卡片的属性在3024个选择集、3个患者角色、9种提示改写和9个实验组中被独立随机化,共产生40068个评分响应;性别和族裔通过姓名传递,采用通信审计方法。声誉信号占主导:将评分从3.9提高到4.7会使选择概率增加31.4个百分点(pp),将费用从90美元提高到190美元会使选择概率降低20.0个百分点。人口统计平等被拒绝,但方向与人类审计研究的预测不同:女性姓名的选择概率增加2.5个百分点,西班牙裔、南亚裔和黑人姓名的选择概率比白人姓名高1.3至2.9个百分点,这种倾斜相当于每次就诊7至14美元的费用差异,且无实质内容的首位列出位置价值11美元。然而,模型在其陈述的理由中提及性别或族裔的比例最多为0.03%,在0.39%的试验中弃权(不执行),因此这些影响在模型自身的解释中不可见,依赖模型自我报告的透明度义务无法检测到它们。一个推理模型完全未通过预先设定的可审计性门槛。该固定设计使审计可重复:任何新模型都可根据相同的刺激进行评估,因此重复的行为审计而非自我报告的解释才是适用的监测技术。
英文摘要
Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.
Comments26 pages, 9 figures, 10 tables