arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.25256cs.AI

谁的对齐?比较不同组织决策情境下的大语言模型过程对齐

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

Niklas Weller, Emilio Barkett

更新

AI总结:

本文提出一种决策策略捕获方法测量过程对齐,发现LLM在ECHR第6条决策中过程对齐与输出准确性高度相关,但在德国消费信贷决策中关系消失,揭示了多元对齐挑战。

AI中文摘要:

将AI系统与组织决策对齐通常被框架化为单一目标问题:使模型表现得像组织一样。我们认为这种框架掩盖了更深的多元主义挑战。我们依赖一种决策策略捕获方法来测量过程对齐:LLM是否像组织一样加权信息,而不仅仅是是否得出相同结论。将此方法应用于ECHR第6条决策,过程对齐强烈预测输出准确性(r = 0.85, p < .001),且外部化显著改善了低对齐模型的对齐。将其应用于德国消费信贷决策,这种关系消失(r = 0.15, p = .60):干预产生不一致的效果,且基准编码了潜在歧视性的历史模式。这种对比本身就是一个多元对齐发现:在有争议的领域,高过程对齐既不能通过外部化实现,也不是无条件可取的。仅凭输出一致性无法区分一个模型是内化了组织政策还是仅仅近似其结果;过程级测量是任何多元对齐评估的必要组成部分。

英文摘要:

Steerable pluralism requires a model to faithfully represent one specified perspective. Organizations are a natural setting for this demand, since they deploy LLMs to make decisions that must reflect their own policy. Yet, most existing work fixes that perspective at the level of individuals or demographic groups. We rely on a decision-policy capturing method to measure process alignment in organizational settings, assessing whether an LLM faithfully reproduces the organization's decision policy rather than merely reaching the same conclusions. We find heterogeneity along two axes. Across models, baseline alignment varies strongly and tracks neither pricing nor general benchmark performance. Across organizations, the structure of alignment changes. In ECHR Article 6 decisions, process alignment predicts output accuracy ($r = 0.85$, $p < .001$), and making the organization's past decision policy explicit improves poorly aligned models. In consumer credit decisions, process alignment is low overall but varies more than output accuracy, and the models resist adopting the organization's weighting of protected attributes. Because historical credit decisions encode potentially discriminatory patterns, higher alignment there is not always desirable. Process-level measurement is therefore necessary, and depending on whether the target policy is normatively desirable, the same procedure can calibrate or audit a model. Deciding which policy to align to, and whether higher alignment is feasible or desirable, makes organizational alignment a pluralistic problem in its own right.

补充信息

↑