AI 中文总结
研究大语言模型对相同症状但不同社会经济地位患者的医疗分诊建议,通过三个模型固定症状改变SES信号,发现对明确信号各模型均提高低SES患者急诊转诊率,对隐含信号的敏感性因模型而异,揭示信号明确性梯度并发布相关内容。
AI 中文摘要
我们研究了仅患者社会经济地位(SES)不同时,大语言模型是否会改变对相同症状的医疗分诊建议。使用三个部署层级模型(Gemini 3.5 Flash、Claude Sonnet 4.6、GPT - 5.4 - mini),固定单一神经症状概况,沿两个渠道改变SES信号:明确的(保险状况、职业、住房)和隐含的(美国邮政编码,无其他社会经济信息)。所有三个模型在有明确信号时,对低SES患者提高了急诊室(ER)转诊率(增幅为13 - 50个百分点)。模型陈述的推理在不同条件下临床几乎相同,推理追踪审计无法察觉这种变化。关键的是,对隐含邮政编码信号的敏感性因模型而异:Gemini仅从地理推断SES,在六个美国邮政编码对中,其ER率汇总变化11.4个百分点(p = 1.4e - 7,6/6对方向相同),Claude Sonnet 4.6保持不变(-0.1个百分点),GPT - 5.4 - mini仅显示微小差异且不一致(2.0个百分点,6对中仅2对预测方向正确)。这揭示了信号的明确性梯度:当SES直接陈述时,每个模型都会对其做出反应,但只有Gemini Flash在必须从仅五位数字的代理中推断时才会对其做出反应。我们将此视为特定于模型的差异,而非大小或成本效应。单句系统提示指令减少但未消除这种效应(Gemini低收入和高收入邮政编码之间的差距从11.4降至5.8个百分点)。我们发布了所有代码、提示和原始结果。
英文摘要
We investigate whether large language models alter medical triage recommendations for identical symptoms when only the patient's socioeconomic status (SES) varies. Using three deployment-tier models (Gemini 3.5 Flash, Claude Sonnet 4.6, GPT-5.4-mini), we hold a single neurological symptom profile fixed and vary the SES signal along two channels: explicit (insurance status, occupation, housing) and implicit (a US ZIP code, with no other socioeconomic information). All three models raise their emergency-room (ER) referral rate for lower-SES patients given the explicit signal (spreads of 13-50 percentage points). The effect is in the protective direction: lower-SES patients are sent to the ER more often, not less. The model's stated reasoning stays clinically near-identical across conditions, so the shift is invisible to a reasoning-trace audit. Critically, sensitivity to the implicit ZIP-code signal is model-dependent: Gemini infers SES from geography alone, shifting its ER rate by a pooled 11.4 points across six US ZIP-code pairs (p = 1.4e-7, same direction in 6/6 pairs), while Claude Sonnet 4.6 stays flat (-0.1 points) and GPT-5.4-mini shows only a small difference that is not sign-consistent (2.0 points, predicted direction in just 2 of 6 pairs), neither a reliable ZIP-code effect, despite both responding to the explicit signal. This reveals an explicitness gradient in the signal: every model acts on socioeconomic status when it is stated outright, but only Gemini Flash acts on it when it must be inferred from a proxy as thin as five digits. We read this as a model-specific difference rather than a size or cost effect. A single-sentence system-prompt instruction reduces but does not eliminate the effect (Gemini's gap between low- and high-income ZIPs falls from 11.4 to 5.8 points). We release all code, prompts, and raw results.
Comments8 pages, 4 tables. Code, prompts, and raw results: https://github.com/wongqihan/triagebench