arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03221cs.CLcs.AIcs.CYcs.LGstat.AP

多步骤临床大语言模型智能体的反事实公平性审计需要测量每行动不稳定性基线

Instability Floors: Separating Bias from Noise in Fairness Audits of Clinical LLM Agents with FairMedAgent

Rohith Reddy Bellibatlu, Manpreet Singh, Deepak Parashar, Rahul Joshi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究指出临床LLM智能体反事实审计的翻转率需结合每行动不稳定性基线解读,发现基线异质且无法通过重复运行完全消除,发布了评估工具FairMedAgent。

中文摘要 AI 辅助

反事实审计是检查临床智能体是否对人口统计学特征不同但临床状况相同的患者区别对待的标准工具,它报告翻转率:仅改变患者描述符时,智能体行动发生变化的频率。我们表明,该量本身无法解释。在16个 vignette(相同叙事、相同描述符字符串、无任何变化)上,对相同条件重复运行10次,临床智能体的行动在8.7%的结果-vignette单元中发生变化,且行动间的不稳定性存在8倍异质性:从ICU升级的0.022到管制药物谨慎的0.179。我们数据中的任何人口统计学对比都无法与该基线区分开来。第二个模型给出的合并基线为6.7%,且对6个行动的排名几乎相同(斯皮尔曼相关系数0.94,精确p值=0.017),因此该基线并非单个系统的人为产物。对5次采样的多数投票聚合消除了39%的不稳定性后趋于平稳,零模拟将剩余部分归因于单元间异质率,因此重复运行可缓解但无法消除不稳定性。因此,任何未附带每行动基线的反事实公平性估计都不能作为差异的证据。测量使用FairMedAgent完成,这是一款用于评估临床大语言模型智能体行动差异的评估工具,其估计量为范围内反事实翻转率,仅统计已发布决策规则允许且经临床医生裁定的行动间翻转。该估计量需要分组裁定,目前正在进行中;此处不主张任何差异结果。每个合成 vignette 在固定形式条件下运行6阶段轨迹(围绕确定性环境步骤的5个模型面向决策),条件涵盖种族、性别、年龄、保险、英语水平及其交互。该工具、基线协议及所有分析脚本均已发布。

英文摘要

Counterfactual fairness audits of clinical language-model agents report a flip rate: how often an action changes when only the patient's demographic descriptor changes. Part of that rate is not demographic. A stochastic agent also changes its own action when nothing changes, and a flip rate cannot be interpreted without knowing how often. We measured it. Re-running one condition ten times over sixteen synthetic vignettes at default sampling changed a clinical agent's action in 8.7 percent of replicate pairs, from 2.2 percent for intensive-care escalation to 17.9 percent for controlled-substance caution, an output given no operational criteria. Across six models from five vendors, pooled floors ranged from 2.5 to 23.7 percent; in this panel neither disclosed size, vendor, nor hosting ordered them. The floor depends on the decoding configuration: majority voting over five draws removed 39 percent of it (95 percent confidence interval (CI) 18 to 64); at temperature 0 three of four locally served models showed no disagreement, but a hosted model still did. We also show that, for a binary action, the flip rate expected under no demographic effect equals the floor and a real effect adds only its square, so a flip rate inside the floor is not evidence of fairness, and direction must be tested with a signed paired test. We give a four-step reporting procedure and release FairMedAgent, the harness, with its protocol, vignettes and analysis scripts, so any team can measure the floor for its own agent.

发表机构

  • Florida International University(佛罗里达国际大学)
  • Boston University(波士顿大学)
  • Symbiosis International University
  • Manipal Institute of Technology(曼ipal理工学院)
  • Manipal Academy of Higher Education(曼ipal高等教育学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑