发表机构
Institute for Future Technologies(未来技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过多智能体模拟发现,在LLM资源分配中拆分决策为多智能体流程未显著改变偏差发生率,但审计能力影响偏差发现率,风险排序审核可提升覆盖范围。
AI 中文摘要
先前的基准测试研究表明,单个大语言模型(LLM)在被迫做出关乎生死的资源分配决策时,会表现出可测量的人口统计学偏差。然而,实际部署场景很少使用单个智能体,而是采用带有审核步骤的流程,旨在精准发现这类失败问题。本研究探讨将同一决策分配给角色差异化的多智能体流程(评估、分配、独立审核),而非由单个模型单独做出并审核时,偏差会发生何种变化。我们使用合成灾难分诊模拟器,该模拟器包含临床特征完全相同但仅存在一项人口统计学属性差异的配对案例,在GPT-4o-mini上运行192轮(共2304个已解决的案例配对),对比单智能体对照条件与九智能体流程在三个独立变化的压力维度下的表现。结果发现,两种条件下出现偏差结果的频率无显著可测量差异(6.9% vs. 6.1%,p=0.498)。但我们发现审计能力对偏差是否被发现存在显著影响:30.0%的偏差结果完全未被发现,当审核员过载时这一比例升至43.8%,而当审核员未过载时则降至18.4%。对该效应的分解显示,其几乎完全由覆盖范围(案例是否被审核,在负载下覆盖范围从100.0%降至65.6%,p<0.001)驱动,而非由被审核案例的判断质量下降导致(81.6% vs. 85.7%,p=1.000,方向相反)。后续实验显示,在相同能力约束下,按估计风险而非先到先服务重新排序审核队列,可恢复大部分损失的覆盖范围(从65.6%升至91.7%,p=0.028)。我们讨论了在资源约束下为LLM智能体流程添加独立监督的系统的影响,并如实报告了研究的局限性:仅使用一种模型、样本量适中且未进行对抗性复现。
英文摘要
Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measurable demographic bias. Real deployments, however, rarely use a single agent: they use pipelines, with review steps meant to catch exactly this kind of failure. We study what happens to bias when the same decision is distributed across a role-differentiated multi-agent pipeline (assessment, allocation, independent audit) instead of made and checked by one model alone. Using a synthetic disaster-triage simulator with paired cases that are clinically identical except for one demographic attribute, we run 192 episodes (2,304 resolved case pairs) on GPT-4o-mini comparing a single-agent control condition to a nine-agent pipeline under three independently varied pressure dimensions. We find no measurable difference in how often biased outcomes occur between the two conditions (6.9% vs. 6.1%, p = 0.498). We do find a large and significant effect of audit capacity on whether bias is caught: 30.0% of biased outcomes go entirely undetected, rising to 43.8% when the auditor is overloaded and falling to 18.4% when it is not. Decomposing this effect shows it is driven almost entirely by coverage (whether a case is reviewed at all, which collapses from 100.0% to 65.6% under load, p < 0.001) rather than by degraded judgment on the cases that are reviewed (81.6% vs. 85.7%, p = 1.000, direction reversed). A follow-up experiment shows that reordering the audit queue by estimated risk, rather than first-come-first-served, recovers most of the lost coverage under the same capacity constraint (65.6% to 91.7%, p = 0.028). We discuss the implications for any system that adds independent oversight to an LLM agent pipeline under resource constraints, and report the study's limitations honestly: one model, modest sample sizes, and no adversarial replication.
Comments6 pages, 2 figures, 3 tables. Code and data available at https://github.com/Polpii/policy-town