大语言模型决策中的代理依赖未与预测证据校准
Proxy reliance in large language model decisions is uncalibrated to predictive evidence
- University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对LLM决策中代理依赖未与预测证据校准的问题,通过临床排序任务测量四种LLM的因果代理效应,发现依赖程度跟不上证据、社会标签抑制脆弱,且准确率评估无法检测这些问题。
AI中文摘要:
大语言模型(LLMs)正进入分诊和信贷发放等决策场景,在此类场景中必须区分与任务相关的推理和不允许的代理使用。当前的审计方法会询问决策是否会随人口统计特征的变化而改变,但与受保护群体相关的属性具有预测价值,因此决策改变可能是歧视也可能是合理推理。我们在一个具有已知真实值的临床排序任务中,测量了四种LLM的因果代理效应,其中证据应有的依赖程度可精确计算并作为参考。一种审计信号得出三种结论:过度依赖、合理依赖和依赖不足。在中性标签下,所有模型都依赖于无信息的代理;在有信息的代理下,所有三种结论均存在。社会领域名称会使依赖程度下降,其中一个模型的依赖程度低于参考值。两项发现解释了这一现象:依赖程度严重跟不上证据,且社会标签抑制作用脆弱,因为上下文示例会使所有模型的依赖程度升至零以上;基于准确率的评估无法检测到上述任何情况。
英文摘要:
Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.